Appearance
Read Aloud
Read Aloud lets you listen to articles and other text sources in ForeverLM. It is designed for the moments when reading is awkward but listening works: walks, chores, commutes, or reviewing a source before a Feynman session.
Start listening
Open a source in the Reader and choose Read Aloud from the toolbar or source menu. ForeverLM extracts the readable text, prepares audio, and shows playback controls at the bottom of the reader. To prepare the complete audio without starting playback, right-click the source in Studio and choose Generate Audio.
While cloud audio is being generated, the progress status names the exact model handling that generation alongside its completion percentage. It remains tied to the in-flight generation even if the model selected in Settings changes before the remaining chunks finish. Web Studio shows the same model and percentage while it generates each chunk.
Sources that already carry their own audio — podcast episodes, YouTube videos, and meeting sessions whose recording plays in the reader — do not offer Read Aloud. Play the original recording instead; generating speech from their transcripts would only duplicate it. Meeting sessions without a playable recording (for example Read AI imports, which are transcript-only) keep Read Aloud.
Playback controls include:
- Pause and resume
- Back 15 seconds
- Forward 15 seconds
- Speed selection from 0.75x to 2x
- Stop
Text-to-speech models
Choose the managed Read Aloud model in Settings -> Text-to-Speech. The catalog includes two free best-effort routes, the economy Kokoro model, and production voices from Fish Audio, Microsoft, Qwen, Mistral, Deepgram, Google, xAI, Sesame, and MiniMax. Managed requests use the ForeverLM AI balance; Grok uses the first-party xAI route and the remaining cloud models use OpenRouter. Voice pickers prefix provider-documented regional accents with their country flag and use 🇺🇳 for multilingual voices instead of a generic globe glyph. The native apps prepare the short “Mitochondria are much more than just the powerhouse of the cell” sample for every model and voice in the background. The currently selected model's visible voices receive a separate high-priority, bounded warm-up instead of waiting behind the catalog-wide queue, so preview buttons play from a durable device cache once preparation completes. Thumbs-up and thumbs-down reactions are saved per model and voice and follow the account across the Mac, iPhone, and Web Studio.
On-device macOS speech and a local Qwen3-TTS server remain free alternatives that do not send narration text to a cloud speech provider. See AI Providers & Models for the complete model list.
Sources that are not in English
Every model in the catalog declares which languages it narrates, and ForeverLM checks a source against that declaration before it generates anything. The check is on-device and deterministic: it samples several windows spread across the source — not just the opening, which is often an English title or abstract — and classifies them. Short, mixed, or ambiguous text is left alone rather than guessed at.
The default, Gemini 3.1 Flash TTS, is the broadest entry in the catalog: Google publishes more than 70 languages for it, so most non-English sources are already covered by the default selection. Most other entries are English-only, and a few name a short list — MAI Voice 2 ships one voice each for English, Spanish, French, and German.
When the source is in a language your selection does not narrate, ForeverLM moves it rather than generating anyway. A sibling voice of the same model wins first (German text goes to MAI's German voice, at the same price and latency), then the default model, then a model with a voice built for that language. The reader tells you which voice it used and which one it replaced. This matters because a model handed a language it does not speak still returns audio and still bills for it — the words are simply read with the wrong pronunciation.
Web Studio does not run this check; it narrates with the model and voice you selected.
Pricing before generation
Read Aloud settings show each cloud model's estimated price per 1 million narration characters and label the Fish Audio S2.1 Pro Free and Deepgram Flux routes as free. Before an uncached source is generated manually, ForeverLM shows its narration length, how long the finished audio runs, and its estimated cost. Grok uses xAI's first-party $15-per-million-character route and is never silently replaced by a second paid model. This confirmation is on by default. Selecting Don't ask again for this model suppresses it only for that model, across all of the model's voices; choosing another model asks again. The matching Read Aloud setting can re-enable confirmation for the selected model.
The estimate uses the model rate available to that app build. The exact managed charge appears in Settings -> Usage after generation; retries can add cost. Replaying cached audio does not generate or charge for new speech.
Automatic audio
Read Aloud can prepare audio on its own, chosen per source type: Articles, Papers, Books, and Courses, in any combination — just articles, or articles and papers, or all four. Every switch starts off. Podcasts, videos, and meetings with a recording have no switch: they arrive as audio already, so ForeverLM never narrates their transcripts. A book's switch covers its chapters, and a course's covers its documents, which is what the counts below say.
The whole feature lives in a collapsed Automatic Audio section, marked as able to spend money, and every switch starts off. One switch can narrate an entire library, so opening the section is deliberate rather than something scrolled past.
A free voice needs no further permission. A paid voice is refused until Allow paid voices is granted, and granting it opens a confirmation that prices the pending library first: every waiting source type with its count, narration characters, listening length, and dollar cost, plus a combined total, at the currently selected voice's rate. Cancel is the default action. A voice with no exact catalog price fails closed no matter what — consent covers a number the learner was shown, never an unknown one — and a type whose backlog cannot be sized is reported rather than folded into a total that would read as complete.
The consent is one account-level setting, not one per source type, and it syncs, so granting it on the Mac authorizes the iPhone too. Each device still spends at whatever rate its own read-aloud selection carries, because that selection is per device. Revoking it stops any paid run immediately and turns off the type switches it was serving. The Mac and iPhone generation workers re-read both the selection and the consent between every source, so a withdrawn consent or a changed voice stops the next item rather than the pass as a whole.
Each type also accepts an optional natural-language rule, such as “papers about mitochondria, and papers where Martin Picard is a coauthor.” Check evaluates the draft against source metadata using the configurable Automatic Audio Filtering background task. It shows every existing source that would be narrated plus the combined narration-character count, listening length, and speech price. The rule is saved only after a successful check; an invalid or failed AI response never broadens the queue. An empty rule means every pending source of that type.
Switching a type on starts every matching source already in the library as well as later arrivals. The Mac, iPhone, and Web Studio preview the sources of that type that have no audio yet, and confirmation is still required because a large free backlog consumes storage and takes time. Saved rules are re-evaluated when relevant library or setting events schedule an automatic-audio pass. A filter failure stops that pass and surfaces the error instead of falling back to all sources.
Once a type is on, background jobs run without further prompts. While a backlog is active, Read Aloud settings show a source-level progress bar with the completed and total counts plus an icon, source type, and title for the item currently being generated, and a Stop button beside it. Stop aborts the run in flight on that device — the deliberate escape hatch for a confirmation clicked by accident. It does not switch the types off or revoke the consent, and it says so: the next source that arrives starts generating again, and whatever was already narrated is billed. Switching the types off is what stops it for good. A source the phone cannot narrate yet — text that has not synced from the Mac, most often — is reported and skipped rather than blocking the sources behind it, and switching the type off and on again retries it.
The iPhone's Automatically continue books setting is separate: it generates the next chapter when the current one finishes, whether or not the Books switch is on.
The switches, rules, and paid consent are all per account and sync, so a type configured from the Mac, the iPhone, or Web Studio behaves the same everywhere. Automatic speech on a free voice never charges the ForeverLM AI balance; on a consented paid voice every generated source is billed to it. A non-empty semantic rule also uses the configurable background text model when it is checked or re-evaluated, so that smaller filtering call has an AI cost of its own regardless of which voice is selected.
Web Studio has the same collapsed section, consent, and confirmation, but it does not generate automatic audio itself — the Mac and the iPhone do — so it has no Stop button.
Listening to a book
Listening to a book chapter runs on into the next chapter when it ends, so starting chapter one commits to the whole book rather than to that chapter. Read Aloud prices it that way: the confirmation names the book and shows two lines, the chapter you started and the whole book, each with its character count, listening length, and estimated cost. Chapters whose audio is already cached are left out of the book total, since replaying them costs nothing.
Approving that price approves the book. Playback moves from chapter to chapter without asking again. Changing the Read Aloud model or voice still invalidates mismatched cached audio, but a model whose confirmation was dismissed stays dismissed when only its voice changes; choosing another unsuppressed model asks again.
The chapter line is measured from the exact narration text. The rest of the book is estimated from each chapter's stored length before its narration is prepared, so the book total is close rather than exact, and a chapter with no stored text yet is left out instead of guessed at.
Web Studio shows the same two lines when you generate audio for a book chapter there, estimating sibling chapters from their word counts. The iPhone shows the book's price in the confirmation itself, because chapter continuation there generates the next chapter in the background without a second prompt: approving one chapter is approving the book, and the alert says so.
Caching
For cloud TTS providers, generated audio can be cached so replaying the same source does not regenerate every chunk. Studio can show a Read Aloud cache indicator for sources with cached audio.
If the source changes or the chosen provider, model, or voice changes, ForeverLM may need to generate fresh audio.
When cloud TTS finishes generating every chunk for a source, ForeverLM writes a synced audio manifest, source text, and chunk files to iCloud Drive. The iPhone can play that audio immediately. If a paper or article has no audio yet, iPhone and CarPlay voice mode can generate it on demand through managed AI, cache it in the same synced format, and start playback without needing the Mac.
Web Studio keeps audio generated there in its bounded, signed-in browser cache. It resumes from completed chunks after a failed request and plays the cached audio from the Reader's Listen button.
iPhone voice agent
The iPhone listener is designed as a voice-first interface, not a manual playback screen. Its realtime voice model uses MCP-compatible app tools, so the same conversation can find and discuss sources, run a review, and change app state. CarPlay uses this same controller and tool set.
The iPhone screen shows listening status, the available synced sources, and the current source. It does not expose manual playback controls for source selection, seeking, speed changes, or play/pause. Those actions are handled by voice.
The agent can control listening and source audio:
- Start playing a source by title, position, current source, or best available default
- Pause, resume, or stop playback
- Move forward or backward in the current audio
- Change playback speed between 0.75x and 2x
- Tell you what source is currently active
- Search the full synced source catalog and inspect the Schedule
- Generate and cache TTS without starting playback
- Generate or reuse TTS and immediately read a paper or article aloud
- Refresh the synced source, Project, Schedule, and audio catalog
The agent can also operate the library:
- List, create, rename, and delete Projects
- Add or remove Project sources
- Add URLs as sources, or archive and delete existing sources
- Assign or move sources to a Schedule day, remove them, or reopen completed work
- Ask for a separate confirmation before any delete, archive, or removal
The agent can enter source chat and review:
- Start a voice review for the current source or another synced source
- Pause the current TTS audio and prepare the source text for questions
- Treat a spoken question during playback as a request to pause and answer from the current text
- Continue a source-grounded conversation once review mode is active
- End review mode and optionally return to audio playback
Typical flows:
- Start the iPhone app in the car and ask it to play an article.
- While listening, ask to pause so you can talk about the article. ForeverLM pauses playback, loads the synced source text, and waits for your question.
- Ask any question about the text. The source chat answers from the synced source text and speaks the response aloud.
- Continue asking follow-up questions naturally.
- Ask to go back to listening when you are done reviewing.
Voice mode is full-duplex and supports barge-in. It requires Sign in with Apple and uses the realtime provider selected in Settings. Source operations are saved through the same CloudKit records the Mac ingests, preserving Project and Schedule relationships across devices.
Good use cases
Read Aloud pairs especially well with Schedule:
- Listen to today's assigned article
- Pause to make a quick note or ask Chat a question
- Start a Feynman review after listening
- Mark the scheduled source complete after the review is linked