July 29, 2026
Grok 4.6 targets Aug 7 at 1.5T, Grok 4.7 follows at 2.1T
Subscribe
Musk's Jul 28 X post puts Grok 4.6 around Aug 7 at 1.5 trillion parameters with upgraded SFT and RL, then Grok 4.7 at 2.1T weeks later with no pricing for either; DeepSeek's officially retired aliases are routing to V4-Flash again per a Jul 28 live probe, but the changelog still has no July entry on day 10; and Kimi K3's technical report lands with Arena AI naming it the top open-weight model for agentic tasks.
Grok 4.6 targets Aug 7, Grok 4.7 weeks later at 2.1T
Elon Musk used an X post on July 28 to set the next two Grok release windows, putting SpaceXAI on a three-model summer after Grok 4.5 shipped July 8.
Grok 4.6 is targeted "around August 7" as a 1.5-trillion-parameter model with "significantly improved SFT and RL," meaning supervised fine-tuning and reinforcement learning. That is the same parameter scale as Grok 4.5's V9 foundation model, so 4.6 looks like a post-training upgrade on the current base rather than a bigger model. Musk disclosed no price, context window, or benchmark numbers for it.
Grok 4.7 follows "a few weeks later" at 2.1 trillion parameters. Musk's framing: "This will be better than 4.6 in every way, except slightly slower to serve, albeit with even better token efficiency." At 2.1T total parameters, Grok 4.7 would trail only Kimi K3 (2.8T) and Qwen 3.8-Max (2.4T) among public models, and would edge past DeepSeek V4 Pro (1.6T). No pricing or availability details were given for either model, and Musk timelines are approximate ("around" Aug 7).
The cadence matters for the pricing beat. Grok 4.5 launched at $2 per million input and $6 per million output with a 500K context window, undercutting Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30) by a wide margin. If 4.6 holds near that rate, SpaceXAI keeps the price-pressure position it has held since July 8. If 4.7's larger 2.1T footprint pushes serving cost up, the token-efficiency claim is the hedge. Vercel CEO Guillermo Rauch recently called Grok 4.5 the top cybersecurity model for price-performance. SpaceXAI (formerly xAI) has not published an official blog post or docs page for 4.6 or 4.7; the timeline is a CEO social post, not a spec sheet.
DeepSeek's retired aliases are routing to V4-Flash again, but the changelog has not moved
The aliases that DeepSeek officially retired on July 24 at 15:59 UTC appear to be live again, at least for some accounts, even as the official changelog stays frozen.
A third-party API tracker ran bounded live probes on July 28 at 22:44 UTC and found that deepseek-chat, deepseek-reasoner, and even the long-retired deepseek-coder all returned HTTP 200 and routed to deepseek-v4-flash. That contradicts the same tracker's July 25 probe, when all three returned HTTP 400. A GET /models call on July 28 still listed only deepseek-v4-flash and deepseek-v4-pro, so the old names are not in the official inventory but are being silently accepted. The tracker's conclusion is careful: this is "observed compatibility on one account, not a documented support guarantee," and "not a reversal of retirement" but a case where "current runtime compatibility and the announced lifecycle do not presently align."
The official DeepSeek changelog, fetched this morning, still shows its latest entry as 2026-04-24 (the V4 Preview launch). There is no July entry, no GA announcement, and no retirement event logged. Day 10 of the "as early as Monday July 20" window that never officially opened. The surge-pricing tier (peak 2x during 9-12 and 14-18 Beijing time) remains announced but not active, and the pricing page still lists a single flat rate: V4 Pro at $0.435 input and $0.87 output per million tokens, V4 Flash at $0.14 and $0.28, both with a 1M context window and 384K max output.
The practical takeaway for anyone who migrated: keep using the explicit V4 IDs. Do not treat the old names as a safe fallback, because the compatibility routing is unannounced and could change without notice. The deeper oddity is that DeepSeek retired the aliases on schedule, never posted the GA to its changelog, and now appears to be quietly re-routing the dead names back to Flash without saying so.
Kimi K3's technical report ships, and Arena AI puts it first among open weights
Moonshot's architecture and evaluation details are now public, and an independent arena has the 2.8T model at the top of the open-weight heap.
The Kimi K3 technical report (submitted July 27) fills in the architecture behind the weights that dropped the same day. Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model that activates 104 billion parameters per token by routing to 16 of 896 experts. It spans 93 layers with a 160K-token vocabulary, uses a 401-million-parameter MoonViT-V2 vision encoder for native image support, and combines 69 Kimi Delta Attention layers with 24 Gated Mixture-of-Latents layers for its hybrid attention system. Moonshot claims 2.5x better scaling efficiency over Kimi K2, attributing the gains to KDA, attention residuals, a sparser MoE configuration, and revised training recipes (though the recipes themselves stay closed).
On benchmarks, Arena AI named Kimi K3 (Max) the number-one open-weight model in its Agent Arena with a 9.75% net-improvement score, surpassing the previous leader GLM-5.2 (Max) at 7.12%, and first place across five evaluation signals. It also took the top open-weight spot in the Frontend Code Arena (1,682 points) and the Text Arena (1,485 points). On the Artificial Analysis Intelligence Index, K3 sits at 57, fourth overall behind Opus 5 (61), Fable 5 (60), and GPT-5.6 Sol (59), but the highest open-weight entry. The hosted API holds at $3 input and $15 output per million tokens with a $0.30 cache hit, flat across the full 1M context.
For self-hosting, the published SGLang deployment matrix specifies eight B300 or MI350X-class GPUs, 16 B200 or H200 GPUs, or 32 80GB H100 GPUs depending on the platform, with some configurations still under final serving verification. A TokenSpeed recipe loads K3 across eight B300 GPUs with about 73GB of memory remaining per GPU after loading. The full Hugging Face repository occupies 1.56 TB. Moonshot recommends the vLLM, SGLang, and TokenSpeed inference engines.

Tracking
- Opus 4.1 retires Aug 5 (7 days). Migrate to Opus 5 at the same $5/$25 rate. Opus 4.1 runs $15/$75. (Anthropic models page)
- Sonnet 5 intro pricing ends Sep 1. $2/$10 through Aug 31, then $3/$15. The new tokenizer produces roughly 30% more tokens, so effective cost from Sep 1 is closer to $3.90/$19.50. (Anthropic pricing)
- Gemini 3.5 Pro still in partner testing. No new date. Google says "as soon as it's ready," Bloomberg reports coding benchmarks fell short after a late-June training-data update. Gemini 4 pretraining has started. (Google blog Jul 21)
- MAI-Cyber-1-Flash public preview Aug 3 in Microsoft Defender and Azure AI Foundry. Covered Jul 28. (SiliconANGLE)
- Ant Ling-3.0-Flash free on OpenRouter through Aug 3, open weights promised after the free period.
- Mistral frontier MoE early access, GA later this summer. No name, specs, or price yet.
- Qwen 3.8-Max-Preview credits-only via Token Plan, no per-token API price. 2.4T params, open weights "soon."
Current prices
Per million input and output tokens, linked to each vendor's official pricing page. DeepSeek verified at off-peak baseline this run (surge tier announced, not active). Rates dated Jul 29 unless noted.
- Claude Fable 5: $10 in / $50 out. Anthropic pricing (verified Jul 26)
- GPT-5.6 Sol: $5 in / $30 out. 1.05M ctx. OpenAI pricing (verified Jul 23)
- Claude Opus 5: $5 in / $25 out. 1M ctx. Anthropic pricing (verified Jul 26)
- Kimi K3: $3 in / $15 out, cache hit $0.30, flat 1M ctx. Kimi blog (open weights live Jul 27)
- Claude Sonnet 5: $2 in / $10 out intro thru Aug 31, then $3/$15. Anthropic pricing
- Gemini 3.6 Flash: $1.50 in / $7.50 out. Google pricing (verified Jul 22)
- Grok 4.5: $2 in / $6 out, 500K ctx. xAI pricing (verified Jul 23)
- Meta Muse Spark 1.1: $1.25 in / $4.25 out, 1M ctx. Meta Model API
- Gemini 3.5 Flash-Lite: $0.30 in / $2.50 out. Google pricing
- Mistral Large 3: $0.50 in / $1.50 out, Apache 2.0. Mistral pricing
- DeepSeek V4 Pro: $0.435 in / $0.87 out, 1M ctx. DeepSeek pricing (verified Jul 29, off-peak)
- DeepSeek V4 Flash: $0.14 in / $0.28 out. DeepSeek pricing (verified Jul 29, off-peak)
Output-price spread: 178x, from Fable 5 at $50 to V4 Flash at $0.28 per million tokens.
That’s the reading for this issue.
Want the next one?
Every new AI Releases & Pricing issue by email. One tap to unsubscribe.