July 22, 2026
Gemini 3.6 Flash Ships at $1.50/$7.50, Undercutting 3.5 Flash as Pro Stalls
Subscribe
Google's stopgap Flash tier landed yesterday while 3.5 Pro stays in partner testing and Gemini 4 pretraining starts; DeepSeek V4 still has not launched 48 hours before its legacy aliases retire; OpenAI retires 15 model snapshots tomorrow; Poolside open-sourced a $0.10/$0.20 coding model.
Google shipped Gemini 3.6 Flash and a cheaper Flash-Lite, but 3.5 Pro is still not ready
The query this beat gets searched on most for Google right now is "Gemini 3.5 Pro release date," and the honest answer is the same one it has been for a week: it is not out. What Google did ship yesterday (July 21) is the stopgap it has been telegraphing, and the pricing is the story.
Google's blog post introduced three Flash-tier models. Gemini 3.6 Flash is the new workhorse, priced at $1.50 per million input and $7.50 per million output, below the $1.50/$9.00 of the 3.5 Flash it succeeds. Google says it consumes 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to 65% fewer on DeepSWE, so the effective cost-per-task cut is larger than the headline rate drop. Benchmarks Google published: DeepSWE 49% vs 37%, MLE Bench 63.9% vs 49.7%, and OSWorld-Verified 83% vs 78.4%, with computer use now a built-in client-side tool. The official pricing page lists batch at $0.75/$3.75 and a Priority tier at $2.70/$13.50.
Gemini 3.5 Flash-Lite is the cheaper, faster sibling at $0.30/$2.50 per million, running at 350 output tokens per second. Google says it beats the older 3.1 Flash-Lite and even outperforms the larger 3 Flash on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%); it is rolling out inside Google Search. Note the older 3.1 Flash-Lite remains the absolute cheapest Google option at $0.25/$1.50 but is roughly 2x slower, per VentureBeat's comparison.
Gemini 3.5 Flash Cyber is fine-tuned on 3.5 Flash for finding and patching security vulnerabilities, but Google is gating it to governments and trusted partners through its CodeMender agent under a limited-access pilot, citing the dual-use risk. No public per-token price was given. MarkTechPost reports it found 55 unique V8 issues versus 47 and 36 for 3.5 Flash and Opus 4.6.
The Pro update, buried at the bottom of the post, is what makes the Flash drop matter strategically: "Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready," adding "we look forward to releasing 3.5 Pro soon." Google also said it has "started our most ambitious pre-training run yet, for Gemini 4." That is the clearest signal yet that the flagship is on an indefinite timeline and the Flash tier is what Google is shipping in the gap, the same stopgap pattern this feed flagged when the 3.6 Flash and 3.5 Flash-Lite names were registered earlier this month. NPowerUser reported on July 22 that Pro is "not ready to go out today," and ETTelecom reported Google released the trio of cheaper models with no timing update for the flagship Pro. Treat any "Gemini 3.5 Pro launched" post as false unless gemini-3.5-pro appears in the public API docs; it still does not.

DeepSeek V4 still has not launched, and the legacy aliases die in 48 hours
The second day of the week brought no DeepSeek V4 general-availability launch either. Primary-source re-verification this morning: the official API Change Log still shows its latest entry as 2026-04-24 (the V4 preview), with no July entry and no GA announcement. The official pricing page still lists only the baseline off-peak rates, V4 Pro at $0.435/$0.87 and V4 Flash at $0.14/$0.28 per million, with cache hits at $0.003625/$0.0028, 1M context and 384K max output, and no peak-pricing row. That is day three of the "as early as Monday" miss that 36kr and The Standard HK reported on July 19, which was always press reporting and never an official DeepSeek statement. (A "deepseek.ai/blog" page claiming a July 24 GA is a fan-site write-up, not the vendor's own docs; the changelog is the proof.)
What is official, and now urgent, is the hard deadline: deepseek-chat and deepseek-reasoner retire on July 24 at 15:59 UTC, which is roughly 48 hours from delivery. Until then those aliases silently route to V4-Flash (non-thinking and thinking respectively, not Pro), so code still calling the old names is already running on Flash. The migration is a rename: switch to deepseek-v4-pro or deepseek-v4-flash. When GA does land, the peak/off-peak billing goes live at 2x the off-peak rate during 9am-12pm and 2pm-6pm Beijing time (01:00-04:00 and 06:00-10:00 UTC), which puts V4 Pro peak at $0.87/$1.74 and Flash at $0.28/$0.56, with the Americas peak windows falling overnight. DeepSeek will be the first frontier API to formalize time-of-day pricing.
OpenAI retires 15 model snapshots tomorrow, July 23
Tomorrow is the OpenAI deprecation wave this feed has been counting down. The official deprecations page lists 15 snapshots shutting down on 2026-07-23 under the April 22 "legacy GPT model snapshots" announcement. The notable ones and their replacements: gpt-5-chat-latest and gpt-5.1-chat-latest route to gpt-5.5; the Codex variants gpt-5-codex, gpt-5.1-codex, gpt-5.1-codex-max, and gpt-5.2-codex go to gpt-5.5 (with gpt-5.1-codex-mini to gpt-5.4-mini); computer-use-preview-2025-03-11 to computer-use-preview or gpt-5.4-mini; the gpt-4o-mini-search-preview and gpt-4o-search-preview dated snapshots to gpt-5.4-mini; gpt-4o-mini-tts-2025-03-20 to the December snapshot; and the deep-research dated snapshots o3-deep-research-2025-06-26 and o4-mini-deep-research-2025-06-26 to gpt-5.5-pro.
The base gpt-5.4 model is NOT on the July 23 list, contrary to the blog post that circulated earlier this month. Separately, the newer July 20 announcement (covered yesterday) retires nine floating-alias audio and realtime families to gpt-realtime-2.1, gpt-audio-1.5, and the mini equivalents on January 20, 2027, so there are two waves in flight from the same vendor this week.
Poolside open-sourced Laguna S 2.1, a coding model at $0.10/$0.20
Also shipping yesterday: Poolside released Laguna S 2.1, a 118-billion-parameter Mixture-of-Experts model that activates 8 billion per token, with a 1M-token context window in both thinking and no-thinking modes. The weights are on Hugging Face immediately under the permissive OpenMDW-1.1 license, and Poolside says it went from the start of training to launch in under nine weeks. The pricing is the eye-catching part: on OpenRouter, VentureBeat reports, Poolside is offering a free 256K-context endpoint and a dedicated 1M deployment at $0.10 per million input and $0.20 per million output, undercutting most frontier options by roughly an order of magnitude.
On the coding benchmarks Poolside published (vendor numbers, no independent verification yet), max thinking lifts Terminal-Bench 2.1 from 60.4% to 70.2% and DeepSWE from 16.5% to 40.4%, and the company published full evaluation trajectories at trajectories.poolside.ai. It is a coding-specialized model from a San Francisco lab that has mostly sold to governments and defense agencies, not a general frontier competitor, so treat the benchmark claims as vendor until tested independently. But at those prices with open weights and a 1M window, it is a notable new entry on the open-weights price ladder.
Tracking
- Kimi K3 open weights, July 27 (5 days). Moonshot is scheduled to release the Kimi K3 weights under its modified-MIT pattern, which would put the 2.8-trillion-parameter model, the largest open-weight ever, on Hugging Face. API pricing stays $3/$15 per million, flat across the full 1M context. Watch for an open-weights leaderboard reshuffle.
- Claude Opus 5 / Honeycomb, still a rumor. As of July 21, Aireiter's tracker confirms Anthropic has announced no Opus 5, no date, and no specs; Opus 4.8 ($5/$25, released May 28) remains the Opus flagship. The community's "July 20-23" launch window is now mostly gone with no launch, the Cursor "Honeycomb EAP" sighting on July 9 remains the only real artifact, and the Vertex AI listing from July 14 is unconfirmed. Label it rumor until an Anthropic model card appears.
- Sonnet 5, September 1. Intro pricing $2/$10 ends August 31, then $3/$15, and the new tokenizer adds roughly 30% more tokens, for an effective $3.90/$19.50, more than Sonnet 4.6's $3/$15.
- Mistral frontier MoE early access is open, GA later this summer, no name, specs, or price yet.
- Fable 5 permanent tier split has been in effect since July 20: Max and Team Premium keep it bundled at 50% of weekly limits, Pro and Team Standard get a one-time $100 credit then $10/$50 per million.
- Qwen 3.8-Max-Preview (launched July 19) remains credits-only on Alibaba's Token Plan with no per-token API price yet and no independent benchmarks.
Current prices
Standard $/1M input/output for the models people actually compare, each linked to the official pricing page. Prices crawled or re-verified July 22, 2026. DeepSeek figures are the off-peak baseline; peak rates (2x) are not yet live.
- GPT-5.6 Sol $5/$30, 1.05M ctx, developers.openai.com
- GPT-5.6 Terra $2.50/$15, 1.05M ctx, developers.openai.com
- GPT-5.6 Luna $1/$6, 1.05M ctx, developers.openai.com
- Claude Fable 5 $10/$50 (cache hits $1, 5m $12.50, 1h $20), platform.claude.com
- Claude Opus 4.8 $5/$25, platform.claude.com
- Claude Sonnet 5 $2/$10 intro thru Aug 31, then $3/$15 (+30% tokenizer = effective $3.90/$19.50), platform.claude.com
- Gemini 3.6 Flash $1.50/$7.50 (batch $0.75/$3.75), ai.google.dev
- Gemini 3.5 Flash-Lite $0.30/$2.50 (batch $0.15/$1.25), ai.google.dev
- Gemini 3.5 Flash $1.50/$9.00, ai.google.dev
- Grok 4.5 $2/$6, 500K ctx, docs.x.ai
- DeepSeek V4 Pro $0.435/$0.87 (off-peak, no peak row yet), api-docs.deepseek.com
- DeepSeek V4 Flash $0.14/$0.28 (off-peak), api-docs.deepseek.com
- Kimi K3 $3/$15 (cache hit $0.30), 1M ctx flat, VentureBeat (no official Moonshot per-token page)
- Meta Muse Spark 1.1 $1.25/$4.25 (cache $0.15), 1M ctx, dev.meta.ai
- Mistral Large 3 $0.50/$1.50, Apache 2.0, mistral.ai
Output-price spread across this list runs about 178x, from $50 for Fable 5 down to $0.28 for DeepSeek V4 Flash.
That’s the reading for this issue.
- DeepSeek V4 Still Not Launched as July 24 API Retirement Looms Jul 21
- Fable 5 Permanent Split Goes Live July 20: Max Keeps It, Pro Pays $10/$50 Jul 20
- Qwen 3.8-Max-Preview Ships at 2.4T Params, Claims Second Only to Fable 5 Jul 19
- Fable 5 Becomes Permanent on Max and Team Premium July 20, Pro Gets a One-Time $100 Credit Jul 18
Want the next one?
Every new AI Releases & Pricing issue by email. One tap to unsubscribe.