July 21, 2026
DeepSeek V4 Still Not Launched as July 24 API Retirement Looms
Subscribe
The "as early as Monday" window closed with no official launch: the API changelog still shows only the April 24 preview, and the legacy deepseek-chat and deepseek-reasoner aliases die July 24. OpenAI separately deprecated nine legacy audio and realtime model families on July 20, retiring them January 20, 2027.
DeepSeek V4 GA still has not landed in the official docs
The question this beat gets searched on most right now is "DeepSeek V4 release date," and the honest answer has not changed in 24 hours: it has not shipped. Chinese tech press (36kr and The Standard HK, both July 19) reported the full-power release could come "as early as Monday" July 20. Monday came and went. As of this morning, the official DeepSeek API Change Log still lists 2026-04-24 as its latest entry, the V4 Preview announcement. There is no July entry, no GA notice, no new model card.
The Models and Pricing page, crawled again today, confirms the same thing from the billing side. It still shows only the baseline off-peak rates: V4-Pro at $0.435 per million input (cache miss) and $0.87 per million output, V4-Flash at $0.14 / $0.28, cache hits at $0.003625 / $0.0028, 1M context with 384K max output. There is no peak-pricing row. That row is the single clearest signal that GA has landed, because DeepSeek told API users in June that peak-valley billing starts with the official release. It is not there yet.
What IS official, and now urgent, is the hard retirement three days out. Per the same April 24 changelog entry and the pricing page footnote, the legacy aliases deepseek-chat and deepseek-reasoner become inaccessible after July 24, 2026, 15:59 UTC. They currently route to V4-Flash non-thinking and thinking modes respectively, so the model behind them is not changing, only the name. The migration is a rename: deepseek-chat to deepseek-v4-flash (or deepseek-v4-pro for quality), deepseek-reasoner to deepseek-v4-pro in thinking mode. Anything still pointed at the old names on July 25 returns model-not-found.
When GA does land, expect the peak-valley mechanic to switch on at the same moment: peak hours 9:00 to 12:00 and 14:00 to 18:00 Beijing time (01:00 to 04:00 and 06:00 to 10:00 UTC), with every rate doubling. That puts V4-Pro peak at $0.87 / $1.74 and V4-Flash peak at $0.28 / $0.56. For teams in the Americas, both peak windows fall overnight, so the 2x rarely bites during US business hours. DeepSeek's own line is still "mid-July," and mid-July ends this week, so the window is genuinely closing, but treat any aggregator headline claiming "GA launched" as false until the changelog moves.

OpenAI deprecated nine legacy audio and realtime families on July 20
The OpenAI deprecations page carries a new entry dated July 20, 2026: legacy audio, realtime, and transcription model families are deprecated and removed from the API on January 20, 2027. Nine slugs are on the list, and the replacements are the newer 2.1 and 1.5 generation that OpenAI has been quietly moving voice traffic onto.
The retiring families and their recommended replacements: gpt-realtime and gpt-4o-realtime to gpt-realtime-2.1; gpt-realtime-mini and gpt-4o-mini-realtime to gpt-realtime-2.1-mini; gpt-audio, gpt-4o-audio, and gpt-audio-mini to gpt-audio-1.5; gpt-4o-mini-audio to gpt-audio-1.5; and gpt-4o-mini-transcribe-2025-03-20 to gpt-4o-mini-transcribe-2025-12-15. The replacement models themselves shipped earlier this month, gpt-realtime-2.1 and -mini on July 6, with a stated 25 percent p95 latency cut across the Realtime line.
This is the second half of a two-wave cleanup of OpenAI's legacy voice stack, and the first half lands in two days. The July 23 wave (covered below) retires the dated snapshots gpt-audio-mini-2025-10-06 and gpt-realtime-mini-2025-10-06. The July 20 announcement retires the floating aliases and the broader gpt-4o-audio and gpt-4o-realtime families on January 20, 2027, six months out. If your audio or realtime integration still references a gpt-4o-* voice slug by name, January 20, 2027 is the hard cutoff; the safer move is to point at gpt-realtime-2.1 or gpt-audio-1.5 now, since the old slugs are already the deprecated path.
Tracking
OpenAI July 23 deprecation, 2 days out. Fifteen model aliases shut down July 23, 2026, per the deprecations page: the chat aliases gpt-5-chat-latest and gpt-5.1-chat-latest to gpt-5.5; five Codex variants (gpt-5-codex, gpt-5.1-codex, gpt-5.1-codex-max, gpt-5.1-codex-mini, gpt-5.2-codex) to gpt-5.5 or gpt-5.4-mini; computer-use-preview and the two gpt-4o-*-search-preview slugs to gpt-5.4-mini; gpt-4o-mini-tts-2025-03-20 to the December snapshot; the dated gpt-audio-mini-2025-10-06 and gpt-realtime-mini-2025-10-06; and o3-deep-research and o4-mini-deep-research to gpt-5.5-pro. The base gpt-5.4 is NOT on the list, despite recurring blog claims. Gateway configs that hardcode any of these names fail on July 24.
Kimi K3 open weights, 6 days out. Moonshot AI shipped Kimi K3 on July 16 at 2.8 trillion parameters, the largest open-weight model ever, $3 / $15 per million tokens with a flat 1M context. The weights were not released at launch; they are scheduled for July 27. When they land, the open-weight size ranking reshuffles: K3 2.8T ahead of Qwen 3.8 2.4T, DeepSeek V4-Pro 1.6T, Inkling 975B, Mistral Large 3 675B.
Sonnet 5, September 1. Claude Sonnet 5 holds at $2 / $10 through August 31, then moves to $3 / $15 on September 1. The new tokenizer emits roughly 30 percent more tokens than the old one, so the effective rate is closer to $3.90 / $19.50, above Sonnet 4.6's $3 / $15.
Gemini 3.5 Pro, still missing. The Google DeepMind model page still lists 3.5 Pro as "coming soon," and the Gemini API pricing page still has no 3.5-pro row, only 3.5 Flash at $1.50 / $9 and 3.1-pro-preview at $2 / $12. Posts claiming "Gemini 3.5 Pro launched July 17" are the same content-mill false reports flagged last week; Google's spokesperson line to Yahoo Tech remains "currently testing 3.5 Pro, an upgraded Flash model, and other models with partners," with no date.
Fable 5 permanent split, now in effect. The tier split went live July 20: Max and Team Premium keep Fable 5 bundled at 50 percent of weekly limits indefinitely; Pro and Team Standard get a one-time $100 credit then $10 / $50 per million tokens via usage credits. The Claude Code 50 percent weekly boost ended the same night, compounding the Max cut.
Qwen 3.8-Max-Preview, credits-only. Alibaba's 2.4T preview (launched July 19) remains Token Plan credits only with no per-token API price yet; the Qwen dev pricing page still lists no qwen3.8 row. Open weights "soon."
Mistral frontier MoE. Early access only, no name, no specs, no price; general availability later this summer.
Current prices
$/1M input / output, linked to each vendor's official pricing page. DeepSeek re-verified today (still off-peak baseline, no peak row). Others last verified July 18 to 20. Where a tokenizer inflates effective cost, it is noted.
- GPT-5.6 Sol $5 / $30, 1.05M context, 128K max out. developers.openai.com
- GPT-5.6 Terra $2.50 / $15. developers.openai.com
- GPT-5.6 Luna $1 / $6. developers.openai.com
- Claude Opus 4.8 $5 / $25, 1M context. platform.claude.com
- Claude Fable 5 $10 / $50 (cache hit $1, 5m $12.50, 1h $20), now the top GA Claude rate. platform.claude.com
- Claude Sonnet 5 $2 / $10 intro through Aug 31, then $3 / $15 (new tokenizer adds ~30% tokens, effective ~$3.90 / $19.50). platform.claude.com
- Gemini 3.5 Flash $1.50 / $9 (batch $0.75 / $4.50), 1M context. ai.google.dev
- Gemini 3.1 Pro Preview $2 / $12 ($4 / $18 above 200K). ai.google.dev
- Grok 4.5 $2 / $6, 500K context (cache $0.50, $1.00 above 200K). docs.x.ai
- DeepSeek V4 Pro $0.435 / $0.87 off-peak (peak 2x pending, no peak row yet). api-docs.deepseek.com
- DeepSeek V4 Flash $0.14 / $0.28 off-peak. api-docs.deepseek.com
- Mistral Large 3 $0.50 / $1.50, Apache 2.0 open weights, 675B / 41B active. mistral.ai
- Meta Muse Spark 1.1 $1.25 / $4.25 (cache $0.15), 1M context, first paid Meta model. dev.meta.ai
- Kimi K3 $3 / $15 (cache hit $0.30), flat across 1M context, 2.8T params. Moonshot via VentureBeat
The output-price spread across this list is roughly 178x, from $50 per million for Fable 5 down to $0.28 for DeepSeek V4 Flash. Qwen 3.8-Max-Preview is omitted because it has no published per-token API price yet.
That’s the reading for this issue.
- Fable 5 Permanent Split Goes Live July 20: Max Keeps It, Pro Pays $10/$50 Jul 20
- Qwen 3.8-Max-Preview Ships at 2.4T Params, Claims Second Only to Fable 5 Jul 19
- Fable 5 Becomes Permanent on Max and Team Premium July 20, Pro Gets a One-Time $100 Credit Jul 18
- Gemini 3.5 Pro Misses July 17 as Moonshot Ships Kimi K3, the Largest Open-Weight Model Ever Jul 17
Want the next one?
Every new AI Releases & Pricing issue by email. One tap to unsubscribe.