July 19, 2026
Qwen 3.8-Max-Preview Ships at 2.4T Params, Claims Second Only to Fable 5
Subscribe
Alibaba's third Chinese open-weight giant in a month landed credits-only on Token Plan today with no per-token API price and no independent benchmarks, DeepSeek V4's full launch is reported as early as Monday with peak-valley pricing, and Fable 5's permanent tier split goes live tomorrow as the Claude Code 50% boost ends tonight.
Alibaba launches Qwen3.8-Max-Preview, 2.4T params, "second only to Fable 5"
The query people search most on this beat today is "Qwen 3.8", and it is real. Alibaba's Qwen team announced Qwen3.8 on X this morning (July 19, 2026): a 2.4-trillion-parameter model the team calls "one of the most powerful model[s] available today, compatible to leading frontier AI models, second only to Fable 5." The preview, Qwen3.8-Max-Preview, went live immediately on Alibaba's Token Plan, Qoder, and QoderWork, with a free trial on the Qwen desktop client. The formal open-weight release is promised "soon," per Yicai (第一财经, July 19).
Two honest caveats up front. First, the "second only to Fable 5" claim rests entirely on Alibaba's own internal evaluations. No independent benchmark from Artificial Analysis or LMArena has been published yet, and the model has no public model card or API per-token price. As OfficeChai notes, "whether Qwen3.8 holds up once outlets like Artificial Analysis and LMArena run their own numbers will decide how much of this claim survives contact with independent testing." Second, this is a preview that follows Alibaba's recent closed-then-open pattern: Qwen 3.6 Max Preview (April) was the first Qwen Max to ship closed, hosted only on Qwen Studio and Alibaba Cloud, with open weights later. Qwen3.8 takes the same path for now.
Access and pricing, the fine print that matters on this beat. Qwen3.8-Max-Preview is credits-only via Token Plan, not a per-token API. The Qwen Cloud developer pricing page lists per-token rates for qwen3.7-max ($2.50/$7.50), qwen3.6-max-preview ($1.30/$7.80 up to 128K), and qwen3.6-flash ($0.25/$1.50), but no qwen3.8-max-preview row. The Token Plan itself is a credits subscription: Individual Lite (2,500 credits per 7 days), Standard (10,000), and Pro (40,000), with 5-hour caps of 700/3,000/12,000 credits respectively, per the Token Plan docs. Team tiers run Standard Seat $20/mo (limited-time, was $30), Pro Seat $75/mo (was $100), and Max Seat $200/mo. A launch promo drops Qwen3.8-Max-Preview daytime credit consumption to 1折 (90% off), with an extra nighttime discount for Individual plans, per Yicai's Token Plan report. Token Plan works with Claude Code, Cursor, Cline, OpenCode, Codex, Kilo CLI, and OpenClaw through OpenAI- and Anthropic-compatible protocols. Notably, the Token Plan model list also includes deepseek-v4-pro alongside qwen3.8-max-preview, qwen3.7-max, and wan2.7-image, so Alibaba is reselling DeepSeek's frontier model inside its own credits bundle.
Where it sits in the open-weight race. At 2.4T params, Qwen3.8 lands just behind Kimi K3 (2.8T) and ahead of DeepSeek V4 Pro (1.6T) and Thinking Machines' Inkling (975B), making it the second-largest open-weight model disclosed. The scale race among Chinese labs has moved fast enough that "largest" is a short-lived title, and Alibaba owns roughly 36% of Moonshot AI, the lab behind Kimi K3, so Qwen3.8 vs Kimi K3 is partly an in-house rivalry. Three major Chinese open-weight announcements in roughly a month (Z.ai's GLM 5.2, Kimi K3, now Qwen 3.8) have each claimed a spot at or near the top of the global leaderboard. Until independent benchmarks land, treat "second only to Fable 5" as Alibaba's pitch, not a verified result.

DeepSeek V4's full launch reported as early as Monday, peak-valley pricing waiting in the wings
DeepSeek's official V4 launch window is "mid-July," and Chinese tech press now reports the full-power version could land as early as Monday, July 20. 36kr reported July 19 that "the official release of DeepSeek V4 may happen as early as tomorrow, and no later than the next few days," with grayscale-test access already circulating to selected users. The Standard (Hong Kong) carried the same "as early as Monday" framing, adding that performance approaches GPT-5.6. This is reporting, not an official DeepSeek announcement, and DeepSeek has not confirmed a date. The DeepSeek pricing page crawled today still lists only baseline off-peak rates with no peak-pricing row, so the launch has not landed as of this writing.
What is confirmed is the pricing structure that activates at GA, set out in DeepSeek's June 29 announcement and the preview docs. Peak hours run 9:00-12:00 and 14:00-18:00 Beijing time (1-4am and 6-10am UTC), during which token prices double. Off-peak rates stay at the current preview baseline: deepseek-v4-pro $0.435/$0.87 per million input(cache miss)/output, rising to $0.87/$1.74 at peak; deepseek-v4-flash $0.14/$0.28, rising to $0.28/$0.56 at peak. Cache-hit input stays near-free at $0.0028/$0.003625 per million. This is the first time-of-day billing scheme on a frontier API, and the practical play for cost-sensitive teams is to shift batch jobs, benchmarks, and data generation outside the two peak windows.
The hard deadline the reader has been tracking: deepseek-chat and deepseek-reasoner retire July 24, 2026 at 15:59 UTC, five days from today, per the footnote on DeepSeek's own pricing page. The old endpoints currently route to deepseek-v4-flash non-thinking/thinking modes. Any integration still pointing at deepseek-chat or deepseek-reasoner needs to move to deepseek-v4-pro or deepseek-v4-flash before that cutover or it breaks.
Fable 5's permanent tier split goes live tomorrow, and the Claude Code boost ends tonight
The Fable 5 cliff this feed tracked through three wire-moves resolves into policy tomorrow. As covered in yesterday's issue, Anthropic announced July 18 that from July 20 Max and Team Premium keep Fable 5 included at 50% of weekly limits indefinitely, while Pro and Team Standard lose included access and get a one-time $100 usage credit, after which Fable 5 costs $10/$50 per million tokens via usage credits (double Opus 4.8's $5/$25, tied with Mythos 5 at the top of Anthropic's rate card).
What changes tonight: the 50% Claude Code weekly usage-limit boost ends July 19 at 11:59:59pm PT, per the same announcement and Anthropic's Claude Code promotion page. After tonight, Code weekly ceilings revert to standard tier levels, so the 50% Fable cap on Max and Team Premium tomorrow sits on top of a lower base than during the promo. Practical read from TechTimes: the $100 credit buys roughly 2 million output tokens at $50 per million, modest for a single autonomous code migration, so Pro users running long agentic tasks should budget for credits or move routine work to Sonnet 5. Free, standard Enterprise, usage-based Enterprise, and API access are unaffected; the API Fable 5 rate stays $10/$50 with cache hits at $1.
Tracking
- Gemini 3.5 Pro: still missing, no new news. The July 17 target came and went with no launch; Bloomberg reported July 16 that the model is months behind on coding, and a Google spokesperson confirmed to CNBC it is "currently testing 3.5 Pro, an upgraded Flash model, and other models with partners." Alphabet dropped roughly 4% on the report. The Gemini API pricing page still lists only 3.5 Flash at $1.50/$9 and 3.1-pro-preview at $2/$12, with no 3.5 Pro row, model card, or API ID. Treat any "Gemini 3.5 Pro launched" post as false unless
gemini-3.5-proappears in the public API docs. - OpenAI July 23 deprecation wave: 4 days out. Per the OpenAI deprecations page, the July 23 sunset covers gpt-5-chat-latest, gpt-5.1-chat-latest, five Codex variants, computer-use-preview, and deep-research models. The gpt-5.4 base model is NOT on the list.
- Kimi K3 open weights: 8 days out, scheduled July 27. K3 shipped July 16 at 2.8T params and $3/$15 per million tokens with a flat 1M context, per VentureBeat; weights follow July 27, which will reshuffle the open-weight leaderboard.
- Sonnet 5 Sep 1 price jump: intro $2/$10 ends August 31, then $3/$15, and the new tokenizer adds roughly 30% more tokens, making the effective rate about $3.90/$19.50 from September 1, more than Sonnet 4.6's $3/$15, per the Anthropic pricing page.
- Mistral frontier MoE: still in early access, GA later this summer, no name, specs, or price yet, per TechCrunch. Mistral Large 3 ($0.50/$1.50, Apache 2.0, 675B/41B-active) remains the current flagship.
Current prices, July 19
All figures are $ per 1M input / 1M output tokens, linked to each vendor's official pricing or docs page (crawled today unless noted). Cache, batch, and tiered long-context rates vary, follow the link for the fine print.
OpenAI (developers.openai.com): GPT-5.6 Sol $5/$30, Sol Fast $12.50/$75, Terra $2.50/$15, Luna $1/$6. All three tiers share 1.05M context, 128K max output, cache writes at 1.25x with 90% read discount and 30-minute minimum cache life.
Anthropic (platform.claude.com): Claude Fable 5 $10/$50 (cache hits $1, 5m $12.50, 1h $20), Opus 4.8 $5/$25, Sonnet 5 $2/$10 intro through August 31 then $3/$15. Fable 5 is the top of Anthropic's rate card, double Opus 4.8.
Google (ai.google.dev): Gemini 3.5 Flash $1.50/$9 (batch/flex $0.75/$4.50), Flash-Lite $0.25/$1.50, 3.1-pro-preview $2/$12. Still no Gemini 3.5 Pro row.
xAI (docs.x.ai): Grok 4.5 $2/$6, 500K context (half of grok-4.3's 1M). Not on the batch-discount list (grok-4.3 and 4.20 get 20% off); Priority tier is a 2x premium.
DeepSeek (api-docs.deepseek.com): V4 Pro $0.435/$0.87 (cache hit $0.003625), V4 Flash $0.14/$0.28 (cache hit $0.0028), 1M context, 384K max output. Baseline off-peak rates shown; peak hours (9-12 and 14-18 Beijing) double these once the full V4 launch goes live.
Alibaba Qwen (docs.qwencloud.com + Token Plan): Qwen3.8-Max-Preview is Token Plan credits only, no per-token API price yet. Qwen3.7-Max $2.50/$7.50 (limited-time 50% off), Qwen3.6-Max-Preview $1.30/$7.80 (to 128K, then $2.00/$12.00), Qwen3.6-Flash $0.25/$1.50.
Meta (dev.meta.ai): Muse Spark 1.1 $1.25/$4.25, $0.15 cached, 1M context with active management, thinking tokens billed at output rate, web search $2.50 per 1,000 queries. Meta Model API public preview, US-only waitlist.
Mistral (mistral.ai): Mistral Large 3 $0.50/$1.50, Apache 2.0, 675B/41B-active MoE. Frontier open-weight MoE in early access with GA later this summer.
Moonshot Kimi (VentureBeat, no official per-token pricing page crawled): Kimi K3 $3/$15, cache hit $0.30, flat across the full 1M context with no long-context premium, web search $0.015 per call. K2.7 Code and K2.6 $0.95/$4.
Thinking Machines Inkling (tinker-docs): $1.87/$4.68 per million at 64K context, cached at 0.20x ($0.374), $3.74/$9.36 at 256K. 975B/41B-active MoE, Apache 2.0 open weights, 1M context (256K on Tinker). Inkling was spared the July 17 platform-wide Tinker hike, so this launch rate still holds.
Prices verified against each vendor's official page on July 19, 2026, except Kimi K3 (VentureBeat) and Inkling (Tinker docs). If a row disagrees with the linked page, the linked page wins.
That’s the reading for this issue.
- Fable 5 Becomes Permanent on Max and Team Premium July 20, Pro Gets a One-Time $100 Credit Jul 18
- Gemini 3.5 Pro Misses July 17 as Moonshot Ships Kimi K3, the Largest Open-Weight Model Ever Jul 17
- Inkling Is Mira Murati's First Model, Priced at $1.87/$4.68 With a Launch Rate That Expires Tomorrow Jul 16
- Gemini 3.5 Pro's July 17 Target Is Two Days Out and Still Unconfirmed, While DeepSeek V4 Peak Pricing Looms Jul 15
Want the next one?
Every new AI Releases & Pricing issue by email. One tap to unsubscribe.