AI Releases & Pricing

July 27, 2026

Kimi K3 Open Weights Are Live at ~594GB, License Reported Modified MIT

Subscribe
Listen

Moonshot's 2.8-trillion-parameter Kimi K3 weights are downloadable on Hugging Face under a Modified MIT license most reporting confirms, with one guide claiming Apache 2.0 unverified, and self-hosting needs roughly 1.4TB of VRAM; DeepSeek's surge-pricing GA still has no changelog entry eight days past the alias deadline despite a wave of content-mill "GA" posts, and Opus 5 holds at $5/$25.

Kimi K3's open weights are on Hugging Face, but read the license file first

The largest open-weight model ever built is now downloadable. Moonshot AI published Kimi K3's weights to its Hugging Face organization around the July 27 00:00 UTC target it set at launch, with explainx tracking the drop at roughly 7:30 PM EDT on July 26, slightly ahead of schedule, and Crypto Briefing confirming public availability on July 27. The native MXFP4 safetensors release is about 594GB, and the repo is gated, so you accept a license agreement before the download starts.

The model itself is unchanged from the July 16 API launch: 2.8 trillion total parameters on a Mixture-of-Experts design that activates 16 of 896 experts per token (roughly 50 billion active), a one-million-token context window, native vision, and Moonshot's Kimi Delta Attention plus Attention Residuals architecture. Moonshot's own blog is blunt about where it stands: "its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol." On the Artificial Analysis Intelligence Index it scores 57, fourth overall behind Opus 5 (61), Fable 5 (60), and Sol (59), but first among open weights.

The license is the fine print that matters, and it is not fully settled. Crypto Briefing, Glitchwire, and MoClaw (citing Decrypt and Yahoo Finance from the July 18-19 weekend) all name it Modified MIT, matching the K2 line. The K2 precedent, per TechTimes, attaches an attribution requirement (displaying the model name in the product interface) only above 100 million monthly active users or $20 million in monthly revenue. One dev.to self-hosting guide claims Apache 2.0, which conflicts with every other outlet and could not be verified against the gated license file. Treat "Modified MIT" as the reported term, not a contract you have read, until you open the LICENSE file in the repo.

Self-hosting is a datacenter proposition, not a workstation one. The practical Q4 MXFP4 target is about 1.4TB of VRAM, roughly 18 H100 80GB GPUs, at about $50 an hour reserved on AWS, Azure, or GCP. Moonshot recommends 64 or more accelerators for production, and local inference on day one caps context near 131K tokens versus the full 1M on the hosted API. A vLLM build with KDA prefill-cache support ships alongside the weights, so the OpenAI-compatible serving path is ready.

Kimi K3 self-host day one
What it takes to run the 2.8T open weights on your own GPUs. Sources: Moonshot docs, AIToolsRecap, TechTimes (Jul 25-27, 2026)

Managed inference moved on day zero. Together AI launched K3 on July 27 through its Provisioned Throughput offering (guaranteed tokens per minute, 99% uptime SLA), claiming 65% lower cost than Fable 5 on a DeepSWE coding comparison, with serverless API access listed as coming soon. Modal also confirmed day-zero hosted access. The Moonshot API itself stays at $3 per million input tokens and $15 per million output, flat across the full 1M context, with cache hits at $0.30, and it remains free in the consumer Kimi app. The economics to watch are not today's $3/$15, which is an exclusivity-window price, but where Together, Fireworks, and Modal land on identical weights over the next month as the serving market competes.

DeepSeek V4 "GA" headlines are content mills; the official changelog has no July entry

If you have seen posts this week declaring "DeepSeek V4 went generally available on July 19" or "July 20," check them against the primary source. DeepSeek's official API change log, fetched this morning, still lists 2026-04-24 as its latest entry, the V4 preview that introduced deepseek-v4-pro and deepseek-v4-flash. There is no July post, no GA announcement, and no surge-pricing switch-over date. This is the eighth day of the "as early as Monday July 20" window that 36kr and The Standard HK reported on July 19, which was reporting, not an official DeepSeek statement.

The surge pricing is announced, not active. Peak hours of 9-12 and 14-18 Beijing time (01-04 and 06-10 UTC) are supposed to double rates, taking V4 Pro to $0.87/$1.74 and Flash to $0.28/$0.56. But the Deep Seek Fan Hub tracker (an independent guide on a look-alike domain, not official DeepSeek) confirms that as of July 25-26 "the official rate card still lists a single flat tier per model" and the surcharge has "no percentage and no start date." The off-peak baseline is still what the pricing page shows: V4 Pro at $0.435/$0.87, Flash at $0.14/$0.28, 1M context with 384K max output.

The "GA" articles are catching up to a fact, not reporting a launch. Sites like AIToolsReview (July 21) and tech-insider.org (July 24, a known content-mill byline) date a "general availability" to July 19-20, but V4 Pro and Flash have been the production model IDs since the April 24 preview, and the legacy deepseek-chat and deepseek-reasoner aliases retired on schedule on July 24 at 15:59 UTC (confirmed dead by a live API check the next morning). What those posts call "GA" is the existing production availability, not the surge-pricing graduation that DeepSeek itself tied to "mid-July" and has not posted to its changelog. If you are searching "DeepSeek V4 GA," the official change log is the only source that counts, and it has not moved.

Tracking

Current prices, $ per 1M input/output

Official pricing pages, re-verified July 27 unless noted. The output spread from top to bottom is about 178x.

That’s the reading for this issue.