July 23, 2026
OpenAI Retires 15 Models Today and Adds Hard API Spend Caps
Subscribe
The July 23 shutdown wave hits gpt-5-chat-latest, five Codex variants, and deep-research snapshots; OpenAI also shipped hard monthly spend limits that 429 the API on overage. DeepSeek's legacy aliases die tomorrow with V4 GA still not landed; Opus 5 "preparations" reported but unconfirmed.
OpenAI's July 23 deprecation wave goes live today
Today, July 23, 2026, is the shutdown date OpenAI set on April 22 for a 15-model retirement batch. If your code still calls any of these, the endpoint stops responding today, usually with no warning you would catch in advance.
The five Codex coding models are the most likely to break production agents: gpt-5-codex, gpt-5.1-codex, gpt-5.1-codex-max, and gpt-5.2-codex all route to gpt-5.5, while gpt-5.1-codex-mini routes to gpt-5.4-mini. The two chat aliases, gpt-5-chat-latest and gpt-5.1-chat-latest, also move to gpt-5.5. The computer-use and search snapshots, computer-use-preview-2025-03-11, gpt-4o-search-preview-2025-03-11, and gpt-4o-mini-search-preview-2025-03-11, all move to gpt-5.4-mini. The two deep-research models, o3-deep-research-2025-06-26 and o4-mini-deep-research-2025-06-26, move to gpt-5.5-pro. Rounding out the list: gpt-4o-mini-tts-2025-03-20 to gpt-4o-mini-tts-2025-12-15, gpt-audio-mini-2025-10-06 to gpt-audio-1.5, and gpt-realtime-mini-2025-10-06 to gpt-realtime-mini.
Two things worth flagging. First, the base gpt-5.4 model is NOT on this list, it stays live. Second, the current generation is now gpt-5.6 (Sol/Terra/Luna), so if you are migrating anyway it is worth testing against 5.6 rather than only the recommended 5.5 replacement. OpenAI itself notes these migrations are rarely pure drop-ins: output format, tone, and tool-calling behavior can shift, so test before you cut over. The full list and replacements are on the official deprecations page, and a developer community writeup lays out the same 15 entries with migration targets.
OpenAI ships hard spend limits (new, July 22)
The second OpenAI platform change this week is genuinely new and has not shipped before. On July 22 the API changelog added hard spend limits for organizations and projects.
The mechanism, per the spend limits guide: you set a monthly dollar cap at the organization level, the project level, or both. A soft limit is monitoring only. A hard limit makes API responses return a 429 error once tracked spend reaches the monthly cap, and they keep failing until you raise the limit or it resets the next month. Spend alerts fire at thresholds you pick, so you get notified before traffic is actually interrupted.
For a beat that tracks what models cost, this is the first native OpenAI kill switch on overspend rather than just an alert. It is aimed at fixed-budget experiments, dev projects, and customer-specific workloads where a runaway agent loop could otherwise rack up a large bill silently. It does not change any per-token price, but it changes how predictably you can cap what you spend.
DeepSeek V4 GA still not landed, legacy aliases die tomorrow
This is day four of the missed "as early as Monday" window, and the primary sources still show no launch. The official API change log was crawled again this morning: the latest entry remains 2026-04-24 (the V4 preview), with no July entry and no GA announcement. The pricing page still lists only baseline off-peak rates, with no peak-pricing row: V4 Pro at $0.435 input / $0.87 output per million tokens, V4 Flash at $0.14 / $0.28, cache hits at $0.003625 and $0.0028, 1M context with 384K max output.
The hard deadline that does land is tomorrow. At 15:59 UTC on July 24, the legacy aliases deepseek-chat and deepseek-reasoner stop working. Until that moment they silently route to V4 Flash (non-thinking and thinking modes), not to V4 Pro. The migration is a rename: point deepseek-chat workloads at deepseek-v4-flash and deepseek-reasoner workloads at deepseek-v4-flash with thinking on, or move up to deepseek-v4-pro if you want the full-power model.
One debunk worth repeating because it ranks high on a "DeepSeek V4 GA" search. A post at deepseek.ai/blog/deepseek-v4-ga-surge-pricing-migration, authored by "Deep Seek Fan Hub" and dated July 21, claims V4 goes GA on July 24. That domain is a fan site, not DeepSeek, and the claim is not an official announcement. When GA does land, expect peak pricing to switch on: 2x the off-peak rate during 9 to 12 and 14 to 18 Beijing time (01:00 to 04:00 and 06:00 to 10:00 UTC), which for the Americas falls overnight.
Opus 5 / Honeycomb: fresh "preparations" report, still a rumor
The reader has been watching this one, so here is the honest update. Claude Opus 5 has not shipped. Anthropic's current Opus flagship remains Opus 4.8 at $5 / $25 per million tokens, released May 28. There is no Opus 5 entry in Anthropic's news archive or models docs as of today.
What is new since the last run is thin and single-sourced. A tracker writeup updated with a July 23 sighting reports that "on July 23 a model tracker reported fresh launch preparations," citing an @M1Astra post, and notes social posts have moved the next rumored date to July 24. Neither carries an Anthropic source. The concrete artifact behind all of this is still the "Honeycomb EAP" listing that briefly appeared in Cursor around July 8 to 9 (a 1M context window, an "xhigh" reasoning mode, and a safety fallback to Opus 4.8), then was pulled within hours, plus an unconfirmed claude-opus-5 string reportedly seen on Vertex AI around July 14.
Label all of this RUMOR. There is no model card, no API model ID, no pricing, and no Anthropic announcement. The one grounded timing signal is Anthropic's own Opus cadence: 4.6 on February 5, 4.7 on April 16 (a 70-day gap), 4.8 on May 28 (42 days). By July 23 it has been 56 days since 4.8, a stretch that is growing but not yet the longest in the 4.x run, which is part of why "imminent" keeps trending without anyone at Anthropic saying so. The reliable place to watch is Anthropic's news page and models docs, where a launch would appear first. Do not architect production around a leaked codename.
Tracking
- Kimi K3 open weights, July 27 (4 days). Moonshot's 2.8T-parameter model is live on the API at $3 / $15 with a $0.30 cache hit, but the open weights are not on HuggingFace yet. Scheduled release July 27.
- Sonnet 5 price step, September 1. Intro rate $2 / $10 runs through August 31, then $3 / $15. The new tokenizer emits roughly 30% more tokens, so the effective rate from September 1 is about $3.90 / $19.50, more than Sonnet 4.6's $3 / $15. (Anthropic pricing)
- Gemini 3.5 Pro, still missing. Google shipped the 3.6 Flash stopgap on July 21 ($1.50 / $7.50, 17% fewer output tokens) and said 3.5 Pro is "testing with partners" and will be broadly available "as soon as it's ready," with no date. Gemini 4 pretraining has started. Treat any "3.5 Pro launched" post as false unless
gemini-3.5-proappears in the public API docs. - Fable 5 permanent split, in effect since July 20. Max and Team Premium keep Fable 5 at 50% of weekly limits; Pro and Team Standard get a one-time $100 credit then pay $10 / $50 via usage credits.
- Qwen 3.8-Max-Preview, credits-only. Alibaba's 2.4T model is live on Token Plan with no per-token API price yet; open weights "soon."
- Mistral frontier MoE, early access. GA later this summer, still no name, specs, or price.
Current prices (verified July 23)
All prices are per 1 million input / output tokens unless noted, and each links to the vendor's official pricing page crawled today.
- GPT-5.6 Sol: $5 / $30, 1.05M context, 128K max output. Sol Fast $12.50 / $75.
- GPT-5.6 Terra: $2.50 / $15, 1.05M context.
- GPT-5.6 Luna: $1 / $6, 1.05M context.
- Claude Opus 4.8: $5 / $25. Current Opus flagship.
- Claude Fable 5: $10 / $50 (cache hit $1, 5m $12.50, 1h $20). Top of the Claude rate card.
- Claude Sonnet 5: $2 / $10 intro through Aug 31, then $3 / $15. With the +30% tokenizer, effective Sep 1 is about $3.90 / $19.50.
- Gemini 3.6 Flash: $1.50 / $7.50 (batch $0.75 / $3.75). Replaces 3.5 Flash ($1.50 / $9).
- Gemini 3.5 Flash-Lite: $0.30 / $2.50, 350 tok/s.
- Grok 4.5: $2 / $6, 500K context. Not on the batch-discount list.
- DeepSeek V4 Pro: $0.435 / $0.87, off-peak (peak 2x pending GA). 1M context, 384K max output.
- DeepSeek V4 Flash: $0.14 / $0.28, off-peak. Legacy aliases
deepseek-chat/deepseek-reasonerretire tomorrow, July 24 15:59 UTC. - Kimi K3: $3 / $15, cache hit $0.30, flat across full 1M context. Open weights due July 27.
- Meta Muse Spark 1.1: $1.25 / $4.25, $0.15 cached, 1M context. First paid Meta model.
- Mistral Large 3: $0.50 / $1.50, 675B / 41B-active MoE.
The output-price spread across this list is 178x, from Fable 5 at $50 down to DeepSeek V4 Flash at $0.28.
That’s the reading for this issue.
- Gemini 3.6 Flash Ships at $1.50/$7.50, Undercutting 3.5 Flash as Pro Stalls Jul 22
- DeepSeek V4 Still Not Launched as July 24 API Retirement Looms Jul 21
- Fable 5 Permanent Split Goes Live July 20: Max Keeps It, Pro Pays $10/$50 Jul 20
- Qwen 3.8-Max-Preview Ships at 2.4T Params, Claims Second Only to Fable 5 Jul 19
Want the next one?
Every new AI Releases & Pricing issue by email. One tap to unsubscribe.