July 28, 2026
Kimi K3 License is MIT-like, but a $20M Model-as-a-Service clause applies
Subscribe
Moonshot's Kimi K3 License reads like MIT but requires a separate Moonshot deal for Model-as-a-Service operators above $20M and UI attribution above 100M users; Microsoft shipped its first cybersecurity model MAI-Cyber-1-Flash for an Aug 3 preview; DeepSeek V4's surge-pricing GA is still absent from the official changelog on day 9; and Opus 5's max effort costs 94% more than the default high for two index points.
The Kimi K3 license is now readable, and it is not MIT
The one open question from the weights release is settled. The LICENSE file on the Kimi K3 Hugging Face repository is now readable, and it is neither MIT nor Apache 2.0. It is a bespoke document Moonshot calls the Kimi K3 License, tagged license:other on Hugging Face, pulled directly from the repo by analysts who read it.
The correction matters: K2 shipped in July 2025 under what Moonshot called a "modified MIT" license. Simon Willison notes that the K3 license "no longer calls itself 'modified MIT' and goes further," adding a clause K2 never had. Outlets that labeled K3 "Modified MIT" were carrying the K2 framing forward; the actual document is more restrictive.
For most of its length it reads like MIT. It grants anyone, free of charge, the right to use, copy, modify, distribute, sublicense and sell the software, and to run, deploy, fine-tune or build derivative works from the weights. Then two conditions attach, and VentureBeat published the full text:
- Section 2, the new one. A "Model as a Service" business, defined as giving third parties inference or fine-tuning access with meaningful control over inputs, parameters or training data, must sign a separate agreement with Moonshot once the aggregate revenue of the licensee and its affiliates exceeds $20 million over any consecutive 12 months. The threshold is on group revenue, not revenue attributable to K3, so a small subsidiary of a larger parent can be caught. It excludes end-user products that embed the model in a specific feature, and mere relaying of requests to a model someone else hosts.
- Section 3, the attribution clause K2 already had. Any commercial product or service with more than 100 million monthly active users, or more than $20 million in monthly revenue, must display "Kimi K3" prominently in its user interface.
- Section 4 exemptions. Purely internal use, and access through Moonshot's own products or its certified inference partners, are exempt from both.
The practical split is by business model, not by user. A team fine-tuning K3 on its own infrastructure carries no obligation. A cloud provider reselling K3 inference at scale needs a contract with Moonshot first. The dev.to guide that claimed Apache 2.0 yesterday was simply wrong; the file is license:other.
The hosted API is unchanged at $3 per million input tokens and $15 per million output, with cached input at $0.30, flat across the full 1-million-token context. What the weights add is an alternative to that meter. Artificial Analysis scores K3 at 57 on its Intelligence Index, fourth behind Opus 5 (61), Fable 5 (60) and GPT-5.6 Sol (59), with a GDPval-AA v2 score of 1668 Elo, up from 1190 for K2.6, and a cost per agentic task of $0.94, close to Sol and roughly half Opus 4.8. Now that the checkpoint is downloadable, those numbers can be rerun by anyone rather than taken from Moonshot's harness.
The largest open-weight model ever built is "open" with business-model strings. For most teams it is effectively free and open. For anyone planning to resell it as a service past $20 million in group revenue, there is a gate.
Microsoft's first cybersecurity model, MAI-Cyber-1-Flash, enters preview Aug 3
Microsoft introduced its first in-house cybersecurity model on July 27, alongside an agentic system called Project Perception. MAI-Cyber-1-Flash is a compact, code-tuned derivative of Microsoft's MAI-Thinking-1 line, trained on the company's own exploit and remediation records.
The pricing-relevant detail is the routing split. Microsoft said the model carries roughly 90% of the workload inside MDASH, its multi-model vulnerability scanning harness, and routes the hardest 10% to OpenAI's GPT-5.4. That split, Satya Nadella said, "delivers world-class performance at 50% of the cost of leading models," a vendor-supplied claim that frames the economics rather than a published rate. The system is priced on consumption, metered in what Microsoft calls Security Compute Units, not a per-token API price.
On the public CyberGym benchmark, 1,507 vulnerability reproduction tasks, MDASH running on MAI-Cyber-1-Flash scored 95.95%, against roughly 84% for Anthropic's Mythos and 88.45% for MDASH alone. The earlier MDASH build found 16 previously unknown Windows flaws, four of them critical remote code execution bugs, all fixed in May's Patch Tuesday. Project Perception fields red agents that probe, blue agents that investigate, and green agents that write and deploy fixes, with human sign-off on high-impact actions.
Public preview opens August 3, initially inside Microsoft Defender, and MAI-Cyber-1-Flash becomes available through Azure AI Foundry the same day, subject to Microsoft's customer vetting. It is not a frontier text model and there is no public per-token price, but it extends the pattern this feed tracked on July 24, when Microsoft shipped priced in-house image and voice models: OpenAI's largest backer is building a first-party, cost-competing stack for specific modalities and routing to GPT only for the sliver of work its own models cannot handle. Anthropic previewed Mythos in April, Google shipped Gemini 3.5 Flash Cyber on July 21, and Cisco has pushed a security model of its own, so the cyber-specialist tier is consolidating fast.
Opus 5's "lower cost per task" win is over Fable 5, not over Opus 4.8
Anthropic's headline for Opus 5 is "comparable intelligence to Fable 5 at 26% lower Cost per Task," and that is accurate at max effort. The fine print, from Artificial Analysis's own evaluation, is that the comparison is to Fable 5. At max effort Opus 5 costs $2.03 per Intelligence Index task, which is more than Opus 4.8 at $1.80 and Sonnet 5 at $1.53. The cost win is real against Fable 5's $2.75; it is not a win against the cheaper Claude models.
Where the money actually goes is the effort dial. The per-effort breakdown shows the curve: medium scores 56 for $1,114.96, high scores 59 for $1,973.77, xhigh scores 60 for $2,909.91, and max scores 61 for $3,835.51. High is Anthropic's default. Going from the default high to max buys two index points for 94% more spend and 92% more tokens. The largest single score gain is medium to high, three points for 77% more cost.

The unchanged $5/$25 per-token rate hides how verbose max effort is. Opus 5 generated 100 million output tokens on the Intelligence Index at max, against a cross-model median of 63 million, which Artificial Analysis calls "very verbose." A migration gotcha worth flagging: disabling thinking at xhigh or max returns HTTP 400, so you cannot run the top two effort levels without thinking on. The practical read is that high, the default, is the production baseline; max is a two-point gain for nearly double the eval spend, and should be reserved for reruns where a high-effort task already failed.
DeepSeek V4 surge pricing: still not in the changelog, day 9
The official DeepSeek API change log, fetched this morning, still lists 2026-04-24 as its latest entry. There is no July post and no general-availability announcement, nine days past the "as early as Monday July 20" window that Chinese tech press reported. The peak-hour surge pricing, 2x during Beijing business hours, remains announced but not active.
The independent DeepSeek guide that tracks the official rate card, updated July 26, corroborates this: the legacy alias retirement went ahead on July 24, but the surcharge "did not go live with it," with "no percentage and no start date published and the official rate card still lists a single flat tier per model." Off-peak baseline rates are still what you pay: V4 Pro at $0.435/$0.87 and V4 Flash at $0.14/$0.28 per million input/output, 1-million-token context, 384K max output.
The wave of "GA" posts from content mills is catching up to the fact that V4 Pro and V4 Flash have been production model IDs since the April 24 preview and the legacy aliases retired July 24, not reporting a surge-pricing launch. DeepSeek's own news page warns to rely only on official accounts. If you migrated off deepseek-chat, remember v4-flash defaults thinking on, so old non-thinking traffic needs thinking:{type:"disabled"} set explicitly or costs and latency jump silently.
Tracking
- Ant Ling-3.0-Flash (Business Wire, July 27): the official press release confirms the July 23 release and adds that model weights will be open-sourced after the free period ends August 3. Free on OpenRouter and Vercel through August 3; no model card or benchmarks yet.
- Claude Opus 4.1 retires August 5 (8 days). Migrate to Opus 5 at the same $5/$25, per the Anthropic models overview.
- Claude Sonnet 5: intro $2/$10 runs through August 31, then $3/$15 on September 1. The roughly 30% tokenizer inflation makes the effective September rate about $3.90/$19.50.
- Gemini 3.5 Pro: still missing. No update since the July 21 "currently testing with partners" statement. Gemini 3.6 Flash shipped July 21 as the stopgap, and Gemini 4 pretraining has started. Treat any "3.5 Pro launched" post as false unless
gemini-3.5-proappears in the public API. - Fable 5 permanent split in effect since July 20: Max and Team Premium keep Fable 5 at 50% of weekly limits; Pro and Team Standard get a one-time $100 credit, then $10/$50.
- Mistral frontier MoE: early access open, broader GA later this summer, no name, specs or price. Mistral Large 3 at $0.50/$1.50 remains the current flagship.
- Qwen 3.8-Max-Preview: credits-only via Token Plan, no per-token API price published.
Current prices
Per million input/output tokens, each linked to the official pricing page. DeepSeek re-verified today; Opus 5, Sonnet 5 and Gemini re-verified this week; the rest stable from official-page checks earlier this week. A 178x output spread runs from Fable 5 at $50 to DeepSeek V4 Flash at $0.28.
- Claude Fable 5: $10/$50, API rate (platform.claude.com)
- Claude Opus 5: $5/$25, cache hit $0.50, batch $2.50/$12.50 (platform.claude.com)
- Claude Sonnet 5: $2/$10 intro through Aug 31, then $3/$15 (platform.claude.com)
- GPT-5.6 Sol: $5/$30, cache hit $0.50 (developers.openai.com)
- GPT-5.6 Terra: $2.50/$15, cache hit $0.25 (developers.openai.com)
- GPT-5.6 Luna: $1/$6, cache hit $0.10 (developers.openai.com)
- Gemini 3.6 Flash: $1.50/$7.50, batch $0.75/$3.75 (ai.google.dev)
- Gemini 3.5 Flash-Lite: $0.30/$2.50 (ai.google.dev)
- Grok 4.5: $2/$6, 500K context (docs.x.ai)
- Meta Muse Spark 1.1: $1.25/$4.25, cache $0.15 (dev.meta.ai)
- Mistral Large 3: $0.50/$1.50, Apache 2.0 (mistral.ai)
- Kimi K3: $3/$15, cache hit $0.30, flat across 1M context, open weights live (kimi.com)
- DeepSeek V4 Pro: $0.435/$0.87 off-peak, peak 2x announced not active (api-docs.deepseek.com)
- DeepSeek V4 Flash: $0.14/$0.28 off-peak, peak 2x announced not active (api-docs.deepseek.com)
That’s the reading for this issue.
Want the next one?
Every new AI Releases & Pricing issue by email. One tap to unsubscribe.