July 26, 2026
Kimi K3's 2.8T Open Weights Drop by July 27, Self-Host at ~1.4TB
Subscribe
Moonshot's largest-ever open weights are due tomorrow with the license still unconfirmed and self-hosting cleared at roughly 1.4TB of VRAM; DeepSeek's V4 general availability still has no changelog entry a week past the alias deadline, and Opus 5 holds the intelligence lead at $5/$25.
Kimi K3's open weights land by July 27, but the license is still TBD
The single largest open-weight model ever built hits Hugging Face tomorrow. Moonshot AI's official Kimi K3 blog commits to releasing "the full model weights by July 27, 2026," alongside a technical report, roughly ten days after the 2.8-trillion-parameter model went live as an API-only product on July 16. At delivery time the weights are not yet public: Moonshot's Hugging Face organization still tops out at Kimi-K2.7-Code, and there is no K3 repository to download.
What is new since this feed last tracked the countdown is the concrete cost of running it yourself, and the one thing still missing from the announcement. Per Moonshot's platform documentation and advance tracking collected by TechTimes (July 25), Moonshot recommends 64 or more accelerators for serious deployment. The practical target is a Q4 MXFP4 quantization at roughly 1.4TB of VRAM, which AIToolsRecap (July 25) translates to about 18 H100 80GB cards and roughly $50 an hour reserved on AWS, Azure, or GCP. The MXFP4 download itself is expected around 594GB, per Hugging Face advance tracking. One nuance that quietly changes the economics: local inference on day one is expected to cap at roughly 131,072 tokens of context, while Moonshot's hosted API serves the full 1-million-token window, so teams buying hardware for the long context need to verify the local cap lifts before committing.
The license is the open question. Every prior Kimi release (K2, K2.5, K2.6, K2.7 Code) shipped under a Modified MIT license that permits commercial use with attribution, and multiple trackers expect K3 to follow. But the exact terms, including whether the monthly-active-user threshold that K2.7 Code added carries over, are not published until the files land. Treat any deployment commitment made before reading the K3 license file as provisional. The official blog notes Moonshot is "working closely with inference partners and open-source maintainers," and a vLLM release incorporating the model's Kimi Delta Attention prefill-cache support is expected to ship alongside the weights.
The pricing-beat angle is the part most coverage glosses. The API's $3 per million input and $15 per million output (cache hits $0.30) is an exclusivity-window price: until the weights are public, Moonshot's hosted endpoint is the only place to run K3. The moment the files land, third-party inference providers like Together AI and Fireworks are expected to offer it within roughly 48 hours, quantized builds will follow within days, and price competition on identical weights starts grinding the effective cost down. As one analysis puts it, the number to watch is not today's $3/$15 but what the serving market charges for K3 in September. Self-hosting also removes the data-residency question that keeps regulated industries off Moonshot's China-hosted API: a 2.8T open weight running on US or EU cloud GPUs puts no Chinese company in the inference loop.
For context, K3 sits at 57 on the Artificial Analysis Intelligence Index, fourth behind Claude Opus 5 (61), Fable 5 (60), and GPT-5.6 Sol (59), and is the highest-placed open-weight model on the board. It is also the third Chinese open-weight giant in roughly a month, after Z.ai's GLM-5.2 and Alibaba's Qwen 3.8-Max-Preview (2.4T, credits-only, open weights promised).
DeepSeek V4 general availability still has no changelog entry, a week past the deadline
Day seven, and the DeepSeek V4 "full GA" with peak/off-peak surge pricing has still never officially landed. Primary re-verification this run: the official DeepSeek API change log crawled July 26 still shows the 2026-04-24 V4 Preview as its newest entry, with zero July posts. The migration shim came out on time (the legacy deepseek-chat and deepseek-reasoner aliases stopped resolving at 15:59 UTC on July 24), but the promised GA that was supposed to accompany it did not.
The off-peak baseline rates remain the live rates on the pricing page: V4 Pro at $0.435/$0.87 per million input/output, V4 Flash at $0.14/$0.28, 1M context with 384K max out. The peak surge (2x during Beijing business hours, 09:00-12:00 and 14:00-18:00, which is 01:00-04:00 and 06:00-10:00 UTC) is announced but still not confirmed live on the docs. If you have not migrated, the rename is deepseek-chat to deepseek-v4-flash (set thinking explicitly to disabled to preserve old behavior, since V4 defaults thinking on) or deepseek-v4-pro for heavier reasoning. And as this feed has flagged all week: a page at deepseek.ai/blog calling itself "Deep Seek Fan Hub" claims a July 24 GA. It is not an official DeepSeek domain. Verify any GA claim against the change log.
Tracking
- Opus 5, day two: Anthropic's pricing page crawled July 26 confirms Claude Opus 5 at $5/$25 per million input/output, unchanged from Opus 4.8 and half of Fable 5's $10/$50. It is the default on Claude Max, the strongest on Pro, and tops the Artificial Analysis Intelligence Index at 61. Opus 4.1 (claude-opus-4-1) is now listed deprecated at $15/$75 and retires August 5; the migration target is Opus 5 at the same $5/$25 tier. Note the newer tokenizer (4.7 and later) emits roughly 30 percent more tokens than older Claudes for the same text, a fine-print cost factor when comparing against GPT-5.x-normalized prices.
- Gemini 3.5 Pro still missing: Google shipped the 3.6 Flash stopgap on July 21 ($1.50/$7.50) and confirmed Gemini 4 pretraining has started, but 3.5 Pro remains in partner testing with no model card, API ID, or pricing row. Treat any "3.5 Pro launched" post as false unless it appears in the public API docs.
- Sonnet 5 price jump, Sep 1: Introductory $2/$10 runs through August 31, then standard $3/$15. With the roughly 30 percent tokenizer inflation, the effective cost from September 1 is about $3.90/$19.50, more than Sonnet 4.6's $3/$15.
- Fable 5 permanent split in effect since July 20: Max and Team Premium keep Fable 5 bundled at 50 percent of weekly limits indefinitely; Pro and Team Standard get a one-time $100 credit then $10/$50 usage credits. API Fable 5 stays $10/$50.
- Mistral frontier MoE: Early access open, general availability later this summer, still no name, specs, or price. Mistral Large 3 ($0.50/$1.50, Apache 2.0) remains the current public frontier.
- Qwen 3.8-Max-Preview (Alibaba, July 19): 2.4T params, credits-only Token Plan, no per-token API price yet, open weights promised.
- Ant Ling Ling-3.0-flash: Free on OpenRouter through August 3, no public weights or benchmarks at launch.
Current prices, verified July 26
Output USD per 1M tokens across the models people actually compare, each linked to its official pricing page. Opus 5, Sonnet 5, Gemini 3.6 Flash, and the DeepSeek change log were re-verified this run (July 26); the rest are stable from official-page checks earlier this week. The spread from Fable 5 to DeepSeek V4 Flash is 178x on output.

- Claude Fable 5: $10 in / $50 out per MTok, 1M ctx, cache hit $1. Verified Jul 26.
- Claude Opus 5: $5 in / $25 out, 1M ctx, 128K max out, cache hit $0.50, Fast mode $10/$50. Verified Jul 26.
- Claude Sonnet 5: intro $2 in / $10 out through Aug 31, then $3 / $15. 1M ctx. Verified Jul 26.
- GPT-5.6 Sol: $5 in / $30 out, 1.05M ctx, 128K max out. Terra $2.50/$15, Luna $1/$6. Stable this week.
- Gemini 3.6 Flash: $1.50 in / $7.50 out, batch $0.75/$3.75. Verified Jul 26. 3.5 Flash-Lite $0.30/$2.50.
- Grok 4.5: $2 in / $6 out, 500K ctx. Stable this week.
- Kimi K3: $3 in / $15 out, cache hit $0.30, flat across full 1M ctx. Weights due by Jul 27.
- Meta Muse Spark 1.1: $1.25 in / $4.25 out, $0.15 cached, 1M ctx. Stable this week.
- Mistral Large 3: $0.50 in / $1.50 out, Apache 2.0, 675B/41B-active MoE. Stable this week.
- DeepSeek V4 Pro: $0.435 in / $0.87 out (off-peak baseline), 1M ctx, 384K max out. Change log re-verified Jul 26, still Apr 24 entry, peak pricing not yet live.
- DeepSeek V4 Flash: $0.14 in / $0.28 out (off-peak baseline). Verified Jul 26.
That’s the reading for this issue.
Want the next one?
Every new AI Releases & Pricing issue by email. One tap to unsubscribe.