July 30, 2026
Grok Voice Think Fast 2.0 launches at $0.08/min, grok-voice-latest upgrades Aug 5
Subscribe
SpaceXAI's Grok Voice Think Fast 2.0 ships at $0.08 per minute of audio with grok-voice-latest cutting over on August 5, the same day Claude Opus 4.1 retires; OpenAI posts a GPT-5.6 efficiency blog claiming Sol beats Fable 5 on the Artificial Analysis Coding Agent Index at under half the cost with 54% fewer output tokens; and DeepSeek's V4 surge-pricing general availability is still absent from the official changelog on day 11.
Grok Voice Think Fast 2.0 is out, priced per minute, cutting over August 5
SpaceXAI released Grok Voice Think Fast 2.0 on July 29, 2026, its "most capable speech-to-speech voice model," and it is the rare frontier release that comes with a clean, predictable price: $0.08 per minute of audio. That is per-minute billing, not per-token, and the company says it chose the unit so budgets do not move with reasoning length.
On the Artificial Analysis speech-to-speech index SpaceXAI cites, Grok Voice Think Fast 2.0 scores 82.9% overall, ahead of GPT-Realtime-2.1 at 79.1%, the prior Grok Voice Think Fast 1.0 at 75.7%, and Gemini 3.1 Flash at 69.5%. Time to first audio drops from 1.25s to 0.70s, and the model uses roughly 0.4x the reasoning tokens of its predecessor while reasoning in parallel with speech. Transcription accuracy is 1.5 to 2.0x better than Deepgram Nova 3 and ElevenLabs Scribe v2, widening to about 10x in noisy settings.

The fine print is a cutover, not just a launch. On August 5, 2026, grok-voice-latest moves from grok-voice-think-fast-1.0 to grok-voice-think-fast-2.0 automatically, with no prompt edits required. To stay on 1.0, pin grok-voice-think-fast-1.0 before that date. SpaceXAI reports an A/B test on Starlink's voice line (+1 888 GO STARLINK) showed higher sales conversion and support containment. The release sits alongside Grok 4.5 at $2/$6 per million text tokens (500K context) and the Grok 4.6 model targeted around August 7, so SpaceXAI is now shipping on three surfaces at once: a text flagship, a voice line, and a teased 1.5T follow-on.
OpenAI's GPT-5.6 efficiency blog: Sol beats Fable 5 on coding at under half the cost
OpenAI published an engineering post, How GPT-5.6 fuses frontier intelligence with frontier efficiency, on July 29, 2026, restating the family's prices and adding a fresh benchmark claim against Anthropic's top model. With maximum reasoning, GPT-5.6 Sol "outperforms Claude Fable 5 on the Artificial Analysis Coding Agent Index at less than half of the cost," and does it with 54% fewer output tokens, per The New Stack's read of the post. Terra is pitched as GPT-5.5 intelligence at half the price, and Luna as 80% below Sol.
The mechanism is the interesting part. OpenAI says that inside Codex, GPT-5.6 Sol autonomously rewrote and optimized its own production kernels in Triton and Gluon (OpenAI's open GPU programming languages), which cut end-to-end serving costs by about 20%. It also designed and ran hundreds of draft-model architecture experiments and intervened when hardware failed. The figures are OpenAI's own production measurements, not independent evaluation, so treat the coding-agent win and the 20% serving cut as vendor-reported.
No price changes here. The GPT-5.6 card is unchanged: Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per million input/output, all 1.05M context, 128K max output, cache writes 1.25x with 90% read discount and a 30-minute minimum cache life. The value of the post for a pricing beat is the effective-cost angle: if Sol genuinely emits 54% fewer output tokens than Fable 5 on coding agent work, the per-token gap understates the real cost gap, the same way Anthropic's own ~30% tokenizer inflation cuts the other direction on Sonnet 5.
DeepSeek V4 surge-pricing GA: still not in the changelog, day 11
The DeepSeek V4 general-availability launch with peak/off-peak surge pricing, reported "as early as Monday July 20" by 36kr and The Standard HK, still has not landed in the official record. The DeepSeek API change log fetched this morning, July 30, shows its newest entry is still 2026-04-24, the V4 Preview. There is no July entry, no GA announcement, and no posted retirement event for the legacy aliases that were scheduled to die July 24.
The pricing page still lists a single flat off-peak tier: V4 Pro at $0.435/$0.87 and V4 Flash at $0.14/$0.28 per million input/output, 1M context with 384K max output, cache hits at $0.003625/$0.0028. The announced peak 2x multiplier (Beijing 9-12 and 14-18, which is 01:00-04:00 and 06:00-10:00 UTC) has no percentage and no start date active on the docs. An independent guide at deepseek.ai, updated July 26, corroborates: the alias retirement went ahead July 24, but "the peak-hour surcharge did not go live with it," and "the official rate card still lists a single flat tier per model."
The legacy names deepseek-chat and deepseek-reasoner were retired July 24 at 15:59 UTC as scheduled, so the migration forcing function already fired even though the GA never posted. Migrate to deepseek-v4-flash or deepseek-v4-pro explicitly, and set thinking behavior by hand because V4-Flash defaults to thinking on. One fan-site trap worth restating: deepseek.ai/blog is an independent "DeepSeek Fan Hub" guide, not an official DeepSeek channel, and its "GA July 24" claim conflicts with the changelog. DeepSeek's own news page warns to "rely only on our official accounts."
Tracking: two cutovers land August 5, plus the rest of the calendar
August 5 is now a double deadline. The same day grok-voice-latest moves to Grok Voice Think Fast 2.0, Claude Opus 4.1 retires (per the official Anthropic models overview, Legacy accordion). Migrate to Claude Opus 5 at $5/$25, the same price Opus 4.8 charged and half of Fable 5's $10/$50. Three days earlier, on August 3, Microsoft's MAI-Cyber-1-Flash reaches public preview inside Defender and Azure AI Foundry, consumption-priced in "Security Compute Units" rather than per token, and Ant Group's Ling-3.0-Flash ends its free OpenRouter/Vercel period with model weights promised to open-source after.
Looking further out: Grok 4.6 is targeted around August 7 at 1.5T parameters with upgraded SFT and RL, then Grok 4.7 a few weeks later at 2.1T, per Elon Musk's July 28 X post (no pricing for either). Claude Sonnet 5's $2/$10 introductory price ends September 1, rising to $3/$15, and its new tokenizer adds about 30% more tokens, so the effective rate is closer to $3.90/$19.50, above Sonnet 4.6's $3/$15. Gemini 3.5 Pro is still in partner testing with no public model card, API ID, or price; Gemini 3.6 Flash at $1.50/$7.50 remains Google's shipping workhorse. Mistral's frontier "fat but sparse" open-weight MoE is still in early access with no name, specs, or price, while Mistral Large 3 holds at $0.50/$1.50 under Apache 2.0. Qwen 3.8-Max-Preview remains credits-only via Token Plan with no per-token API price.
Current prices
Prices per 1M input/output tokens unless noted, each linked to the official pricing or docs page. DeepSeek and Mistral re-verified July 30; OpenAI, Anthropic, Google, xAI, Meta, Moonshot, Qwen stable from official-page checks this week. Output spread is 178x, from Fable 5's $50 down to DeepSeek V4 Flash's $0.28.
- Claude Fable 5: $10/$50, 1M ctx. API rate; Max + Team Premium get it included at 50% of weekly limits, Pro + Team Standard pay $10/$50 via credits. platform.claude.com (Jul 30)
- Claude Opus 5: $5/$25, 1M ctx (default and max), 128K max out, thinking on by default. Fast mode $10/$50, Claude API only. AA Intelligence Index 61, first overall. platform.claude.com (Jul 30)
- Claude Sonnet 5: $2/$10 intro through Aug 31, then $3/$15, 1M ctx. New tokenizer adds ~30% tokens, effective ~$3.90/$19.50 from Sep 1. platform.claude.com (Jul 30)
- GPT-5.6 Sol: $5/$30, 1.05M ctx, 128K max out. Terra $2.50/$15, Luna $1/$6. Sol Fast $12.50/$75. Cache writes 1.25x, 90% read discount, 30-min min life. developers.openai.com (Jul 30)
- Gemini 3.6 Flash: $1.50/$7.50, 1M ctx, 64K max out. 3.5 Flash-Lite $0.30/$2.50. Batch $0.75/$3.75. ai.google.dev (Jul 30)
- Grok 4.5: $2/$6, 500K ctx. Cached $0.50. Not on the batch-discount list. Grok Voice Think Fast 2.0: $0.08/min of audio. docs.x.ai (Jul 30)
- Mistral Large 3: $0.50/$1.50, 256K ctx, Apache 2.0 open weights, 675B/41B MoE.
mistral-large-latest. mistral.ai (Jul 30, re-verified) - Meta Muse Spark 1.1: $1.25/$4.25, 1M ctx (active management), $0.15 cache. Thinking tokens billed at output rate. dev.meta.ai (Jul 30)
- Kimi K3: $3/$15, cache hit $0.30, 1M ctx flat (no long-ctx premium), web search $0.015/call. Open weights live (~594GB MXFP4) under a Kimi K3 License, MIT-like but with a $20M Model-as-a-Service clause and 100M-MAU UI attribution rule. kimi.com (Jul 30)
- DeepSeek V4 Pro: $0.435/$0.87 (off-peak), cache hit $0.003625, 1M ctx, 384K max out. Peak 2x announced, not active. api-docs.deepseek.com (Jul 30, re-verified)
- DeepSeek V4 Flash: $0.14/$0.28 (off-peak), cache hit $0.0028, 1M ctx, 384K max out. Peak 2x announced, not active. api-docs.deepseek.com (Jul 30, re-verified)
- Qwen 3.8-Max-Preview: credits-only via Token Plan, no per-token API price yet. Qwen 3.7-Max $2.50/$7.50. docs.qwencloud.com (Jul 30)
That’s the reading for this issue.
Want the next one?
Every new AI Releases & Pricing issue by email. One tap to unsubscribe.