July 24, 2026
DeepSeek's deepseek-chat and deepseek-reasoner aliases retire today at 15:59 UTC
Subscribe
The legacy model names stop resolving at 8:59 AM Pacific while V4 general availability still has no entry in the official changelog; Microsoft shipped priced in-house image and voice models, and the Claude Opus 5 rumor gained its first real-name sighting.
The DeepSeek migration deadline lands today, but V4 GA still has not
The deadline this feed has tracked for a week arrives this morning. DeepSeek's two legacy model aliases, deepseek-chat and deepseek-reasoner, are deprecated today, July 24, 2026, at 15:59 UTC (8:59 AM Pacific). After that, any call passing either name returns an error. The fix is a one-string rename: deepseek-chat to deepseek-v4-flash (non-thinking) or deepseek-v4-pro, and deepseek-reasoner to deepseek-v4-flash (thinking mode) or deepseek-v4-pro. Base URL, API keys, and billing are unchanged.
One detail worth re-checking before you cut over: until 15:59 UTC, the aliases route to V4-Flash, not Pro. If your app called deepseek-reasoner assuming it hit the strongest model, you have been on Flash the whole time. Moving to deepseek-v4-pro is an upgrade, so re-test your prompts against Pro rather than assuming parity.
What has not arrived is the general-availability launch that deadline was meant to accompany. DeepSeek's official news page still lists the April 24 V4 Preview as its newest entry, with zero July posts. The "official V4 version" that DeepSeek emailed API users about on June 29 and that Chinese press reported for "as early as Monday" on July 19 has slipped past its mid-July window, with the delay attributed to bundling a first-party coding harness. The deadline ships on time even when the launch does not.
That matters for your bill. The announced peak/off-peak structure, the first time-of-day pricing on a frontier API, doubles rates during Beijing business hours (09:00 to 12:00 and 14:00 to 18:00 Beijing, which is 01:00 to 04:00 and 06:00 to 10:00 UTC). Off-peak stays at the current baseline: V4 Pro $0.435 in / $0.87 out, V4 Flash $0.14 / $0.28 per million tokens. Peak doubles those to $0.87 / $1.74 and $0.28 / $0.56. As of this morning I could not confirm the peak tier is live on the docs, so budget for it but verify the rate card before assuming the surcharge is active.

A reminder on sourcing: a post at deepseek.ai/blog claiming "V4 GA July 24" is a fan site ("Deep Seek Fan Hub"), not DeepSeek's official channel. DeepSeek's own preview post tells readers to "rely only on our official accounts." Treat any GA claim as false unless a July entry appears in the official changelog.
Microsoft ships priced in-house image and voice models, claims up to 89% cost cut versus OpenAI
Microsoft AI pushed two purpose-built models into public preview on July 23, and the pricing is concrete enough to model against current spend.
MAI-Image-2.5-Pro, billed as Microsoft's highest-fidelity image generator, is priced at $5 per 1M text input tokens, $8 per 1M image input tokens, and $106 per 1M image output tokens. That more than doubles the image-output price of the base MAI-Image-2.5 ($47), targeting the premium hero-imagery and in-image-text-rendering tier rather than volume generation. It is available in Microsoft Foundry and the MAI Playground.
MAI-Voice-2-Flash, first shown at Build, is a low-latency text-to-speech model at $15 per 1M characters, 2x faster and 32% cheaper than MAI-Voice-2 (which runs $22 per 1M characters). It is aimed at high-volume call-center and voice-agent workloads.
The deployment list is the pricing-pressure signal. Microsoft says the models already run in Bing, PowerPoint, OneDrive, Dynamics 365, Excel, GitHub Copilot, and Azure, and reports vendor-supplied production figures: up to 84% GPU cost reduction in PowerPoint versus OpenAI's GPT-Image-2, and up to 89% in Dynamics 365 Contact Center (VentureBeat, July 23). Those are Microsoft's own numbers on Microsoft's own workloads, so read them as a vendor claim, not an independent benchmark. But the framing, "Microsoft products, powered by Microsoft models," is the clearest sign yet that OpenAI's largest backer is building a first-party stack that competes with OpenAI's frontier models on cost for specific modalities.
Claude Opus 5 rumor: the codename era may be over, but Anthropic's page still says 4.8
RUMOR. The trail on Claude Opus 5 just produced its most specific artifact yet, and it is still not an announcement.
For weeks every leaked trace used the "Honeycomb" codename. That changed on July 23 to 24: screenshots circulating on X show a Cursor error dialog naming the model outright, "The model claude-opus-5-thinking-high requires Max Mode to be enabled." That is the first sighting of the literal claude-opus-5 string in a shipping product, and the -thinking-high suffix matches how Cursor labels reasoning variants, which reads like launch plumbing rather than an experiment. Deployment tracker @M1Astra posted that Opus 5 had begun rolling out across providers, with some users reportedly served the new model under the "Opus 4.8" label before a full switchover. A separate screenshot, attributed to someone posing as an Anthropic employee, showed a Fable 5 guardrail routing to Opus 5.
Against all of that stands the primary record. Anthropic's official models overview still lists Claude Opus 4.8 (May 28, 2026) as the newest Opus, with no Opus 5 model card, pricing page, or API model ID. A forensic audit published July 24 found no Opus 5 entry in Anthropic's, Google Cloud's, or Cursor's public model catalogs. The community's favorite Thursday target, July 23, came and went with no Anthropic launch post. Until anthropic.com/claude/opus or the API docs ship a named model, the public Opus flagship remains Opus 4.8 at $5 in / $25 out per million tokens. Build on that, watch the rest.
Tracking
- Kimi K3 open weights, July 27 (3 days). Moonshot plans to publish the full 2.8-trillion-parameter weights by July 27, which would be the largest open-weight release ever. The catch is 1.4 TB of weights and a license that has not yet been published, so commercial usability stays unconfirmed until the terms land with the release. The API has served
kimi-k3at $3 / $15 per million tokens since July 16. - Gemini 3.5 Pro still missing. No model card, no API ID, no pricing row on Google's pricing page; the Gemini 3.6 Flash stopgap shipped July 21 and Gemini 4 pretraining has started. Separately, Google's Gemini Spark agent opened to US Google AI Pro subscribers on July 24 ($20/month), with global access gated to Google AI Ultra (~$100).
- Ant Ling Ling-3.0-flash (July 23). InclusionAI, Ant Group's open-source lab, released a 124B MoE model with 5.1B active parameters per token, 256K native context extendable to 1M. It is free on OpenRouter, Novita, Vercel, and ZenMux through August 3; post-promo pricing is not yet published, and no model card, weights, or benchmark table shipped at launch, so the lab's "matches our 1T flagship" claim is currently unverifiable.
- Sonnet 5 price step, September 1. Intro pricing of $2 / $10 per million tokens runs through August 31, then moves to $3 / $15. The new tokenizer emits roughly 30% more tokens, so effective cost from September 1 is closer to $3.90 / $19.50, above Sonnet 4.6's $3 / $15 (Anthropic pricing).
- Claude Opus 4.1 retires August 5.
claude-opus-4-1-20250805is deprecated and retires August 5, 2026, per Anthropic's models page. Migrate to Opus 4.8 before the cutoff. - Fable 5 permanent split is in effect since July 20: Max and Team Premium keep Fable 5 bundled at 50% of weekly limits, while Pro and Team Standard get a one-time $100 credit then $10 / $50 per million tokens via usage credits (Anthropic).
- Mistral frontier MoE remains in early access with no name, specs, or price; general availability is pegged for later this summer.
Current prices
Output and input dollars per 1M tokens, each linked to the official pricing or docs page, verified July 24, 2026. Off-peak unless noted.
- GPT-5.6 Sol $5 in / $30 out, 1.05M ctx. Terra $2.50 / $15, Luna $1 / $6. (developers.openai.com)
- Claude Fable 5 $10 / $50, 1M ctx, cache hit $1, 5m $12.50, 1h $20. (platform.claude.com)
- Claude Opus 4.8 $5 / $25, 1M ctx. (platform.claude.com)
- Claude Sonnet 5 $2 / $10 intro through Aug 31, then $3 / $15. New tokenizer adds ~30% tokens. (platform.claude.com)
- Gemini 3.6 Flash $1.50 / $7.50, batch $0.75 / $3.75. 3.5 Flash-Lite $0.30 / $2.50. 3.5 Flash $1.50 / $9. (ai.google.dev)
- Grok 4.5 $2 / $6, 500K ctx. Not on the batch-discount list; Priority is 2x. (docs.x.ai)
- DeepSeek V4 Pro $0.435 / $0.87 off-peak, peak (announced) $0.87 / $1.74, cache hit $0.003625. V4 Flash $0.14 / $0.28, peak $0.28 / $0.56. 1M ctx, 384K max out. (api-docs.deepseek.com)
- Meta Muse Spark 1.1 $1.25 / $4.25, $0.15 cached, 1M ctx, US-only waitlist. (dev.meta.ai)
- Mistral Large 3 $0.50 / $1.50, Apache 2.0, 675B / 41B-active MoE. (mistral.ai)
- Kimi K3 $3 / $15, cache hit $0.30, flat across full 1M ctx, weights due July 27. (kie.ai)
- Qwen 3.7-Max $2.50 / $7.50. Qwen 3.8-Max-Preview is credits-only on Token Plan with no per-token API price yet. (docs.qwencloud.com)
The output-price spread across this list is roughly 178x, from DeepSeek V4 Flash at $0.28 to Claude Fable 5 at $50 per million tokens.
That’s the reading for this issue.
Want the next one?
Every new AI Releases & Pricing issue by email. One tap to unsubscribe.