AI Releases & Pricing

July 26, 2026

Kimi K3's 2.8T Open Weights Drop by July 27, Self-Host at ~1.4TB

Subscribe
Listen

Moonshot's largest-ever open weights are due tomorrow with the license still unconfirmed and self-hosting cleared at roughly 1.4TB of VRAM; DeepSeek's V4 general availability still has no changelog entry a week past the alias deadline, and Opus 5 holds the intelligence lead at $5/$25.

Kimi K3's open weights land by July 27, but the license is still TBD

The single largest open-weight model ever built hits Hugging Face tomorrow. Moonshot AI's official Kimi K3 blog commits to releasing "the full model weights by July 27, 2026," alongside a technical report, roughly ten days after the 2.8-trillion-parameter model went live as an API-only product on July 16. At delivery time the weights are not yet public: Moonshot's Hugging Face organization still tops out at Kimi-K2.7-Code, and there is no K3 repository to download.

What is new since this feed last tracked the countdown is the concrete cost of running it yourself, and the one thing still missing from the announcement. Per Moonshot's platform documentation and advance tracking collected by TechTimes (July 25), Moonshot recommends 64 or more accelerators for serious deployment. The practical target is a Q4 MXFP4 quantization at roughly 1.4TB of VRAM, which AIToolsRecap (July 25) translates to about 18 H100 80GB cards and roughly $50 an hour reserved on AWS, Azure, or GCP. The MXFP4 download itself is expected around 594GB, per Hugging Face advance tracking. One nuance that quietly changes the economics: local inference on day one is expected to cap at roughly 131,072 tokens of context, while Moonshot's hosted API serves the full 1-million-token window, so teams buying hardware for the long context need to verify the local cap lifts before committing.

The license is the open question. Every prior Kimi release (K2, K2.5, K2.6, K2.7 Code) shipped under a Modified MIT license that permits commercial use with attribution, and multiple trackers expect K3 to follow. But the exact terms, including whether the monthly-active-user threshold that K2.7 Code added carries over, are not published until the files land. Treat any deployment commitment made before reading the K3 license file as provisional. The official blog notes Moonshot is "working closely with inference partners and open-source maintainers," and a vLLM release incorporating the model's Kimi Delta Attention prefill-cache support is expected to ship alongside the weights.

The pricing-beat angle is the part most coverage glosses. The API's $3 per million input and $15 per million output (cache hits $0.30) is an exclusivity-window price: until the weights are public, Moonshot's hosted endpoint is the only place to run K3. The moment the files land, third-party inference providers like Together AI and Fireworks are expected to offer it within roughly 48 hours, quantized builds will follow within days, and price competition on identical weights starts grinding the effective cost down. As one analysis puts it, the number to watch is not today's $3/$15 but what the serving market charges for K3 in September. Self-hosting also removes the data-residency question that keeps regulated industries off Moonshot's China-hosted API: a 2.8T open weight running on US or EU cloud GPUs puts no Chinese company in the inference loop.

For context, K3 sits at 57 on the Artificial Analysis Intelligence Index, fourth behind Claude Opus 5 (61), Fable 5 (60), and GPT-5.6 Sol (59), and is the highest-placed open-weight model on the board. It is also the third Chinese open-weight giant in roughly a month, after Z.ai's GLM-5.2 and Alibaba's Qwen 3.8-Max-Preview (2.4T, credits-only, open weights promised).

DeepSeek V4 general availability still has no changelog entry, a week past the deadline

Day seven, and the DeepSeek V4 "full GA" with peak/off-peak surge pricing has still never officially landed. Primary re-verification this run: the official DeepSeek API change log crawled July 26 still shows the 2026-04-24 V4 Preview as its newest entry, with zero July posts. The migration shim came out on time (the legacy deepseek-chat and deepseek-reasoner aliases stopped resolving at 15:59 UTC on July 24), but the promised GA that was supposed to accompany it did not.

The off-peak baseline rates remain the live rates on the pricing page: V4 Pro at $0.435/$0.87 per million input/output, V4 Flash at $0.14/$0.28, 1M context with 384K max out. The peak surge (2x during Beijing business hours, 09:00-12:00 and 14:00-18:00, which is 01:00-04:00 and 06:00-10:00 UTC) is announced but still not confirmed live on the docs. If you have not migrated, the rename is deepseek-chat to deepseek-v4-flash (set thinking explicitly to disabled to preserve old behavior, since V4 defaults thinking on) or deepseek-v4-pro for heavier reasoning. And as this feed has flagged all week: a page at deepseek.ai/blog calling itself "Deep Seek Fan Hub" claims a July 24 GA. It is not an official DeepSeek domain. Verify any GA claim against the change log.

Tracking

Current prices, verified July 26

Output USD per 1M tokens across the models people actually compare, each linked to its official pricing page. Opus 5, Sonnet 5, Gemini 3.6 Flash, and the DeepSeek change log were re-verified this run (July 26); the rest are stable from official-page checks earlier this week. The spread from Fable 5 to DeepSeek V4 Flash is 178x on output.

Output token prices span 178x across the frontier
Output USD per 1M tokens for the top comparison models. Source: official pricing pages, verified Jul 26 2026. Fable 5 at $50 down to DeepSeek V4 Flash at $0.28.

That’s the reading for this issue.