July 31, 2026
OpenAI slashes GPT-5.6 Luna 80% to $0.20/$1.20, DeepSeek V4 Flash goes official
Subscribe
OpenAI cut GPT-5.6 Luna 80% and Terra 20% on July 30 while adding a 2.5x-speed Sol Fast mode at double the price; DeepSeek answered July 31 with the official V4 Flash release (the changelog's first July entry in 11 days), same $0.14/$0.28 price but stronger agents and native Responses-API support, surge pricing still pending; and Thinking Machines shipped Inkling-Small open weights at a quarter of Inkling's size for $0.58/$1.44.
OpenAI cuts GPT-5.6 Luna 80%, Terra 20%, adds Sol Fast mode
The price war this beat has been tracking since June just got its sharpest move. On July 30, 2026, one day after an engineering blog that restated prices, OpenAI passed its serving-efficiency gains to customers:
- GPT-5.6 Luna, the cost tier, drops 80% to $0.20 / $1.20 per million input/output (from $1 / $6).
- GPT-5.6 Terra, the balanced tier, drops 20% to $2 / $12 (from $2.50 / $15).
- GPT-5.6 Sol is unchanged at $5 / $30, but a new Sol Fast API mode delivers up to 2.5x faster responses at twice the price (effectively $10 / $60), replacing Priority Processing. Existing
priority-tagged requests auto-route to Fast mode. - Sol's reasoning, 1.05M context, and 128K max output are untouched. The models page now lists Terra at $2/$12 and Luna at $0.20/$1.20, dated July 30.
The positioning is blunt. OpenAI quotes Replit's head of AI calling Luna "the closest we've come to intelligence too cheap to meter," and claims Luna beats Fable 5 on Agents' Last Exam at an estimated cost per task ~99% lower. Whether that holds under independent eval, the sticker math is now unambiguous: at $1.20/M output, Luna sits below Google's Gemini 3.5 Flash-Lite ($2.50 combined) and well below Gemini 3.6 Flash ($9 combined), and it undercuts OpenAI's own GPT-5.4 ($2.50/$15) on the work Terra is now meant to absorb. ChatGPT and Codex subscription prices are unchanged; Terra and Luna just consume fewer credits. The cuts began rolling out on AWS the same day.
This is a separate announcement from the July 29 efficiency blog (which carried no price change). It reframes the GPT-5.6 three-tier launch from July 9: Terra was already "GPT-5.5 intelligence at half the cost," and Luna is now priced to compete head-on with the cheapest Chinese open-weights rather than sit above them.

DeepSeek V4 Flash goes official, 11 days after the changelog went quiet
The DeepSeek saga this feed has tracked for eleven days resolved on July 31, but not the way the surge-pricing watch assumed. DeepSeek shipped the official public beta of V4 Flash (build DeepSeek-V4-Flash-0731), confirmed by Nikkei Asia ("made the official beta version of its latest V4 Flash model publicly available on Friday") and visible on the official pricing page: "The deepseek-v4-flash model has been updated to DeepSeek-V4-Flash-0731. The calling method remains unchanged."
What actually changed, and what didn't:
- Same architecture, same price. Still 284B total / 13B-active MoE, 1M context, 384K max output. Pricing is unchanged at $0.14 input / $0.28 output per million (cache hit $0.0028). No increase for the upgrade.
- It is a post-training upgrade, not a new model. DeepSeek reports materially stronger agentic, coding, and tool-calling ability, plus native Responses-API support and specific adaptation for Codex-style coding agents (while still speaking OpenAI ChatCompletions and Anthropic-style interfaces). Vendor benchmarks cited include Terminal-Bench 2.1 at 82.7 and Toolathlon 70.3 (treat as vendor figures).
- Only the Flash API was upgraded. V4 Pro and the consumer app/web are unchanged; DeepSeek says an official V4 Pro release is "coming soon."
- Surge pricing is still pending. Nikkei reports DeepSeek "says it will implement peak-hour pricing plan," language consistent with the announced-not-active status this feed has verified for weeks. The peak 2x windows (Beijing 9-12 and 14-18, i.e. 01-04 and 06-10 UTC) remain a plan, not an invoice line.
For readers who have been watching the changelog stay frozen on the April 24 preview entry: this is the first July movement, and it lands as a Flash-only post-training refresh rather than the full surge-pricing general availability that "mid-July" implied. The legacy deepseek-chat/deepseek-reasoner aliases retired July 24 as scheduled; deepseek-v4-flash now resolves to the -0731 build automatically. The practical read is that DeepSeek chose to ship a stronger cheap model into OpenAI's price-cut window rather than flip on time-of-day billing, which keeps V4 Flash at roughly a third of V4 Pro's $0.435/$0.87 and far below the newly discounted Luna.
Inkling-Small ships open weights at a quarter of Inkling's size
Thinking Machines Lab released Inkling-Small on July 30, two weeks after Inkling, its first model. The pitch is comparable performance at roughly a quarter the size, with VentureBeat reporting it surpasses the larger predecessor on several benchmarks.
Full weights are on Hugging Face, and fine-tuning is available on Tinker. With the limited-time 50% discount, the 64K-context variant runs $0.58 per million prefill (input), $1.44 per million sampled (output), and $1.73 per million training tokens, with cached prefill at $0.116. A 256K-context variant is available at higher rates. It is also chattable on Tinker Playground for text, image, and audio.
This is the second open-weights release from Mira Murati's lab (Inkling itself was 975B/41B-active MoE, Apache 2.0, AA Intelligence Index 41). Inkling-Small fills the efficiency slot the launch post previewed, and it lands inside the same 48-hour window as the OpenAI and DeepSeek moves, adding another data point to the open-weights price ladder at the low end.
Tracking
- Aug 5 is a double cutover, 5 days out.
grok-voice-latestmoves from Grok Voice Think Fast 1.0 to 2.0 automatically (pin 1.0 to stay); and Claude Opus 4.1 (claude-opus-4-1-20250805) retires. Anthropic's deprecations page and models overview now point Opus 4.1 users to Opus 5 ($5/$25, same as Opus 4.8), a drop-in name swap with two breaking changes: thinking is on by default, and disabling thinking atxhigh/maxeffort returns a 400. - Grok 4.6 still targets around Aug 7 at 1.5T params, with Grok 4.7 "a few weeks later" at 2.1T (Musk X post, July 28). No pricing for either.
- MAI-Cyber-1-Flash hits public preview Aug 3 (Microsoft Defender, Azure AI Foundry), consumption-priced in Security Compute Units, not per-token. Ant Ling-3.0-flash stays free through Aug 3, with weights open-sourced after.
- MiniMax H3 launched July 31 as a general-purpose omni-modal generation model (video with native stereo audio up to 2K/15s), with MiniMax promising per-second prices below a third of mainstream models at 2K and weights "in the coming days." It is a video/multimodal generator, not an LLM text tier, so no per-token API price to list yet.
- Sonnet 5 intro pricing $2/$10 runs through August 31, then steps to $3/$15 on September 1, with the new tokenizer adding roughly 30% more tokens (effective ~$3.90/$19.50).
- Gemini 3.5 Pro is still in partner testing; the API still lists only 3.6 Flash ($1.50/$7.50) and 3.5 Flash-Lite ($0.30/$2.50) as the live Google tiers. Gemini 4 pretraining has started.
- Mistral's frontier open-weight MoE remains in early access with no name, specs, or price; Large 3 ($0.50/$1.50, Apache 2.0) is the current public frontier.
- Qwen 3.8-Max-Preview is still credits-only on the Token Plan with no published per-token API price.
Current prices
Prices per million input/output tokens, linked to each vendor's official page and dated. GPT-5.6 Terra and Luna reflect the July 30 cut; DeepSeek V4 Flash reflects the July 31 -0731 build. All other rates stable from official-page checks this week.
- Claude Fable 5: $10 / $50 (platform.claude.com, Jul 31). Cache hit $1.
- GPT-5.6 Sol: $5 / $30, 1.05M ctx (developers.openai.com, Jul 31). Fast mode $10 / $60 (2.5x speed, replaces Priority Processing).
- Claude Opus 5: $5 / $25, 1M ctx (platform.claude.com, Jul 31).
- Kimi K3: $3 / $15, cache hit $0.30, flat 1M ctx (kimi.com, Jul 31). Open weights live, Modified MIT reported.
- Claude Sonnet 5: $2 / $10 intro thru Aug 31, then $3 / $15 (platform.claude.com, Jul 31).
- GPT-5.6 Terra: $2 / $12, 1.05M ctx (developers.openai.com, Jul 31). Was $2.50 / $15.
- Gemini 3.6 Flash: $1.50 / $7.50 (ai.google.dev, Jul 31).
- Grok 4.5: $2 / $6, 500K ctx (docs.x.ai, Jul 31). Grok Voice Think Fast 2.0 $0.08/min.
- Meta Muse Spark 1.1: $1.25 / $4.25, cache $0.15, 1M ctx (dev.meta.ai, Jul 31).
- Mistral Large 3: $0.50 / $1.50, Apache 2.0 (mistral.ai, Jul 31).
- DeepSeek V4 Pro: $0.435 / $0.87, 1M ctx (api-docs.deepseek.com, Jul 31). Peak 2x announced, not active.
- GPT-5.6 Luna: $0.20 / $1.20, 1.05M ctx (developers.openai.com, Jul 31). Was $1 / $6.
- Gemini 3.5 Flash-Lite: $0.30 / $2.50 (ai.google.dev, Jul 31).
- DeepSeek V4 Flash: $0.14 / $0.28, cache hit $0.0028, 1M ctx (api-docs.deepseek.com, Jul 31). Now -0731 build.
Output-price spread: roughly 178x, from Fable 5 at $50 down to DeepSeek V4 Flash at $0.28. Luna's cut narrows the OpenAI-vs-DeepSeek gap but does not close it.
That’s the reading for this issue.
- Grok Voice Think Fast 2.0 launches at $0.08/min, grok-voice-latest upgrades Aug 5 Jul 30
- Grok 4.6 targets Aug 7 at 1.5T, Grok 4.7 follows at 2.1T Jul 29
- Kimi K3 License is MIT-like, but a $20M Model-as-a-Service clause applies Jul 28
- Kimi K3 Open Weights Are Live at ~594GB, License Reported Modified MIT Jul 27
Want the next one?
Every new AI Releases & Pricing issue by email. One tap to unsubscribe.