AI Releases & Pricing

July 31, 2026

OpenAI slashes GPT-5.6 Luna 80% to $0.20/$1.20, DeepSeek V4 Flash goes official

Subscribe
Listen

OpenAI cut GPT-5.6 Luna 80% and Terra 20% on July 30 while adding a 2.5x-speed Sol Fast mode at double the price; DeepSeek answered July 31 with the official V4 Flash release (the changelog's first July entry in 11 days), same $0.14/$0.28 price but stronger agents and native Responses-API support, surge pricing still pending; and Thinking Machines shipped Inkling-Small open weights at a quarter of Inkling's size for $0.58/$1.44.

OpenAI cuts GPT-5.6 Luna 80%, Terra 20%, adds Sol Fast mode

The price war this beat has been tracking since June just got its sharpest move. On July 30, 2026, one day after an engineering blog that restated prices, OpenAI passed its serving-efficiency gains to customers:

The positioning is blunt. OpenAI quotes Replit's head of AI calling Luna "the closest we've come to intelligence too cheap to meter," and claims Luna beats Fable 5 on Agents' Last Exam at an estimated cost per task ~99% lower. Whether that holds under independent eval, the sticker math is now unambiguous: at $1.20/M output, Luna sits below Google's Gemini 3.5 Flash-Lite ($2.50 combined) and well below Gemini 3.6 Flash ($9 combined), and it undercuts OpenAI's own GPT-5.4 ($2.50/$15) on the work Terra is now meant to absorb. ChatGPT and Codex subscription prices are unchanged; Terra and Luna just consume fewer credits. The cuts began rolling out on AWS the same day.

This is a separate announcement from the July 29 efficiency blog (which carried no price change). It reframes the GPT-5.6 three-tier launch from July 9: Terra was already "GPT-5.5 intelligence at half the cost," and Luna is now priced to compete head-on with the cheapest Chinese open-weights rather than sit above them.

Output price per 1M tokens across 12 models, July 31 2026
Output $/1M per model, July 31 2026. Luna was $6 and Terra was $15 before the July 30 cut. Source: official vendor pricing pages.

DeepSeek V4 Flash goes official, 11 days after the changelog went quiet

The DeepSeek saga this feed has tracked for eleven days resolved on July 31, but not the way the surge-pricing watch assumed. DeepSeek shipped the official public beta of V4 Flash (build DeepSeek-V4-Flash-0731), confirmed by Nikkei Asia ("made the official beta version of its latest V4 Flash model publicly available on Friday") and visible on the official pricing page: "The deepseek-v4-flash model has been updated to DeepSeek-V4-Flash-0731. The calling method remains unchanged."

What actually changed, and what didn't:

For readers who have been watching the changelog stay frozen on the April 24 preview entry: this is the first July movement, and it lands as a Flash-only post-training refresh rather than the full surge-pricing general availability that "mid-July" implied. The legacy deepseek-chat/deepseek-reasoner aliases retired July 24 as scheduled; deepseek-v4-flash now resolves to the -0731 build automatically. The practical read is that DeepSeek chose to ship a stronger cheap model into OpenAI's price-cut window rather than flip on time-of-day billing, which keeps V4 Flash at roughly a third of V4 Pro's $0.435/$0.87 and far below the newly discounted Luna.

Inkling-Small ships open weights at a quarter of Inkling's size

Thinking Machines Lab released Inkling-Small on July 30, two weeks after Inkling, its first model. The pitch is comparable performance at roughly a quarter the size, with VentureBeat reporting it surpasses the larger predecessor on several benchmarks.

Full weights are on Hugging Face, and fine-tuning is available on Tinker. With the limited-time 50% discount, the 64K-context variant runs $0.58 per million prefill (input), $1.44 per million sampled (output), and $1.73 per million training tokens, with cached prefill at $0.116. A 256K-context variant is available at higher rates. It is also chattable on Tinker Playground for text, image, and audio.

This is the second open-weights release from Mira Murati's lab (Inkling itself was 975B/41B-active MoE, Apache 2.0, AA Intelligence Index 41). Inkling-Small fills the efficiency slot the launch post previewed, and it lands inside the same 48-hour window as the OpenAI and DeepSeek moves, adding another data point to the open-weights price ladder at the low end.

Tracking

Current prices

Prices per million input/output tokens, linked to each vendor's official page and dated. GPT-5.6 Terra and Luna reflect the July 30 cut; DeepSeek V4 Flash reflects the July 31 -0731 build. All other rates stable from official-page checks this week.

Output-price spread: roughly 178x, from Fable 5 at $50 down to DeepSeek V4 Flash at $0.28. Luna's cut narrows the OpenAI-vs-DeepSeek gap but does not close it.

That’s the reading for this issue.