Appearance
AI Providers & Models
There are no API keys to manage. You pick a provider and model for each AI capability in Settings, and ForeverLM debits the complete selected-upstream cost from your prepaid dollar balance. Sub-cent requests remain precisely metered. Balance additions start at $25, larger additions discount the standard checkout margin, and new accounts start with $1.
Anthropic, OpenAI, Google AI Studio, Meta Model API, and xAI text models connect through their first-party APIs behind ForeverLM's metered proxy. OpenRouter serves the long tail. A single Use Ramp Router switch sends every selected model through Ramp instead. Provider keys stay on the server, and a request never falls back silently to another route.
You can also skip the built-in AI entirely and drive ForeverLM from Claude, Codex, Gemini, or Grok, connected with its tailored setup prompt at 1¢ per tool call. Sources, Schedule, and semantic search work either way. Most people mix the two.
Chat and reviews (text models)
In-app Chat, Feynman reviews, and background tasks (like paper metadata extraction) use the text provider and model selected in Settings → AI Provider. The default is Anthropic with Claude Sonnet 5.
| Provider | Models | Default |
|---|---|---|
| Anthropic | Claude Sonnet 5, Claude Fable 5, Claude Opus 5, Claude Haiku 4.5 | Claude Sonnet 5 |
| OpenAI | GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna | GPT-5.6 Sol |
| Google (Gemini) | Gemini 3.5 Flash, Gemini 3 Pro Preview, Gemini 2.5 Flash Lite | Gemini 3.5 Flash |
| xAI (Grok) | Grok 4.3, Grok 4.5 | Grok 4.3 |
| Meta | Muse Spark 1.2, Muse Spark 1.2 Contributor | Muse Spark 1.2 |
| OpenRouter | DeepSeek V4 Flash, DeepSeek V4 Pro, Kimi K3, GPT OSS 120B, GPT OSS 20B, Qwen3 32B, Qwen3 235B, GLM 5.2, Mistral Large 3, Mistral Small 4, Nemotron 3 Ultra, Nemotron 3 Super, Nemotron 3 Nano | DeepSeek V4 Flash |
Notes:
- All long-tail models — DeepSeek, Kimi, Qwen, GLM, GPT OSS, Mistral, Llama, and Nemotron — are served through OpenRouter while the Ramp switch is off. The picker groups them by the lab that trained them, and you can add any other compatible model from OpenRouter's directory.
- DeepSeek and Kimi are offered only through OpenRouter's US-routed inference; the labs' own APIs are deliberately not integrated.
- Ramp Router is an explicit account-wide transport override. Models stay listed once; enabling Use Ramp Router narrows the model pickers to the intersection with Ramp's authenticated, active, fully priced catalog and sends those selections through Ramp's Responses-compatible API. A previously saved model that is no longer available remains the saved choice but cannot be newly selected or silently replaced.
- First-party models are admitted only while ForeverLM has a complete price for every billable usage dimension. An unknown model or premium processing tier is rejected instead of being estimated.
- The model picker shows a curated list; you can hide models you never use in Settings → AI Models → Chat.
- Each provider remembers its own model choice, so switching providers doesn't reset your selection.
Read Aloud (text-to-speech)
Read Aloud uses one managed speech catalog in macOS, iOS, and Web Studio settings. Gemini 3.1 Flash TTS with Kore remains the default, and it is also the widest-reaching entry: Google publishes more than 70 languages for it, from Afrikaans and Arabic through Mandarin, Swahili, Tamil, and Vietnamese. Every other catalog entry declares a narrower set, and most declare English only.
| Tier | Models | List price |
|---|---|---|
| Free, best effort | Fish Audio S2.1 Pro Free, Deepgram Flux TTS Free | Free |
| Economy | Kokoro 82M | $0.62/M characters |
| Value | Sesame CSM 1B | $7/M characters |
| Production | Fish Audio S2.1 Pro, Grok Voice TTS 1.0, MAI Voice 2 Flash, Qwen Audio 3 TTS Flash | $15/M characters |
| Production | Voxtral Mini TTS, MAI Voice 2, Deepgram Aura 2 | $16–$30/M characters |
| Premium | Gemini 3.1 Flash TTS, MiniMax Speech 2.8 Turbo/HD | about $40–$100/M characters |
The free OpenRouter routes have no production latency or availability guarantee. Fish Audio S2.1 is billed by UTF-8 input bytes upstream, so its character-based picker estimate can be lower than the exact charge for non-ASCII text. Grok uses xAI's first-party managed route; every other entry uses OpenRouter. The exact server debit, rather than the picker estimate, is the billing record.
On-device macOS speech and a local Qwen3-TTS server remain available without cloud inference cost.
Voice input (speech-to-text)
Voice dictation in Chat uses the provider selected in Settings → Speech-to-Text:
- Groq (default): Whisper Large v3 Turbo
- OpenAI: GPT-4o mini Transcribe
Audio is sent only to your configured provider and is not persisted after transcription.
Semantic search
Semantic source search runs entirely on-device using Apple's NaturalLanguage contextual embeddings. It costs nothing and sends nothing over the network.
Costs
- Every provider above is available immediately — nothing is behind a tier, and there is no key to obtain. Signing in with your Apple account is the only setup.
- Settings → AI Usage tracks every request against its complete selected-upstream cost. The dollar balance stays precise below a cent. Models without a published rate are shown as unpriced rather than guessed.
- Your balance carries over — it doesn't expire or reset monthly. Top up when you run out.
- Background tasks (paper metadata extraction, transcript cleanup, search embeddings) draw on the same balance and are itemised per task in the ledger. Each one's model is configurable in Settings → AI Models → Background, so you can point the noisy ones at something cheap.
- If a provider or model can't complete a request, ForeverLM surfaces the error — it never silently falls back to a different model or to heuristic grading.