AI Releases & Pricing

July 28, 2026

Kimi K3 License is MIT-like, but a $20M Model-as-a-Service clause applies

Subscribe
Listen

Moonshot's Kimi K3 License reads like MIT but requires a separate Moonshot deal for Model-as-a-Service operators above $20M and UI attribution above 100M users; Microsoft shipped its first cybersecurity model MAI-Cyber-1-Flash for an Aug 3 preview; DeepSeek V4's surge-pricing GA is still absent from the official changelog on day 9; and Opus 5's max effort costs 94% more than the default high for two index points.

The Kimi K3 license is now readable, and it is not MIT

The one open question from the weights release is settled. The LICENSE file on the Kimi K3 Hugging Face repository is now readable, and it is neither MIT nor Apache 2.0. It is a bespoke document Moonshot calls the Kimi K3 License, tagged license:other on Hugging Face, pulled directly from the repo by analysts who read it.

The correction matters: K2 shipped in July 2025 under what Moonshot called a "modified MIT" license. Simon Willison notes that the K3 license "no longer calls itself 'modified MIT' and goes further," adding a clause K2 never had. Outlets that labeled K3 "Modified MIT" were carrying the K2 framing forward; the actual document is more restrictive.

For most of its length it reads like MIT. It grants anyone, free of charge, the right to use, copy, modify, distribute, sublicense and sell the software, and to run, deploy, fine-tune or build derivative works from the weights. Then two conditions attach, and VentureBeat published the full text:

The practical split is by business model, not by user. A team fine-tuning K3 on its own infrastructure carries no obligation. A cloud provider reselling K3 inference at scale needs a contract with Moonshot first. The dev.to guide that claimed Apache 2.0 yesterday was simply wrong; the file is license:other.

The hosted API is unchanged at $3 per million input tokens and $15 per million output, with cached input at $0.30, flat across the full 1-million-token context. What the weights add is an alternative to that meter. Artificial Analysis scores K3 at 57 on its Intelligence Index, fourth behind Opus 5 (61), Fable 5 (60) and GPT-5.6 Sol (59), with a GDPval-AA v2 score of 1668 Elo, up from 1190 for K2.6, and a cost per agentic task of $0.94, close to Sol and roughly half Opus 4.8. Now that the checkpoint is downloadable, those numbers can be rerun by anyone rather than taken from Moonshot's harness.

The largest open-weight model ever built is "open" with business-model strings. For most teams it is effectively free and open. For anyone planning to resell it as a service past $20 million in group revenue, there is a gate.

Microsoft's first cybersecurity model, MAI-Cyber-1-Flash, enters preview Aug 3

Microsoft introduced its first in-house cybersecurity model on July 27, alongside an agentic system called Project Perception. MAI-Cyber-1-Flash is a compact, code-tuned derivative of Microsoft's MAI-Thinking-1 line, trained on the company's own exploit and remediation records.

The pricing-relevant detail is the routing split. Microsoft said the model carries roughly 90% of the workload inside MDASH, its multi-model vulnerability scanning harness, and routes the hardest 10% to OpenAI's GPT-5.4. That split, Satya Nadella said, "delivers world-class performance at 50% of the cost of leading models," a vendor-supplied claim that frames the economics rather than a published rate. The system is priced on consumption, metered in what Microsoft calls Security Compute Units, not a per-token API price.

On the public CyberGym benchmark, 1,507 vulnerability reproduction tasks, MDASH running on MAI-Cyber-1-Flash scored 95.95%, against roughly 84% for Anthropic's Mythos and 88.45% for MDASH alone. The earlier MDASH build found 16 previously unknown Windows flaws, four of them critical remote code execution bugs, all fixed in May's Patch Tuesday. Project Perception fields red agents that probe, blue agents that investigate, and green agents that write and deploy fixes, with human sign-off on high-impact actions.

Public preview opens August 3, initially inside Microsoft Defender, and MAI-Cyber-1-Flash becomes available through Azure AI Foundry the same day, subject to Microsoft's customer vetting. It is not a frontier text model and there is no public per-token price, but it extends the pattern this feed tracked on July 24, when Microsoft shipped priced in-house image and voice models: OpenAI's largest backer is building a first-party, cost-competing stack for specific modalities and routing to GPT only for the sliver of work its own models cannot handle. Anthropic previewed Mythos in April, Google shipped Gemini 3.5 Flash Cyber on July 21, and Cisco has pushed a security model of its own, so the cyber-specialist tier is consolidating fast.

Opus 5's "lower cost per task" win is over Fable 5, not over Opus 4.8

Anthropic's headline for Opus 5 is "comparable intelligence to Fable 5 at 26% lower Cost per Task," and that is accurate at max effort. The fine print, from Artificial Analysis's own evaluation, is that the comparison is to Fable 5. At max effort Opus 5 costs $2.03 per Intelligence Index task, which is more than Opus 4.8 at $1.80 and Sonnet 5 at $1.53. The cost win is real against Fable 5's $2.75; it is not a win against the cheaper Claude models.

Where the money actually goes is the effort dial. The per-effort breakdown shows the curve: medium scores 56 for $1,114.96, high scores 59 for $1,973.77, xhigh scores 60 for $2,909.91, and max scores 61 for $3,835.51. High is Anthropic's default. Going from the default high to max buys two index points for 94% more spend and 92% more tokens. The largest single score gain is medium to high, three points for 77% more cost.

Opus 5 max effort costs 94% more than the default for 2 index points
Cost to run the Artificial Analysis Intelligence Index by reasoning effort. High is the Anthropic default. Low cost not published. Source: Artificial Analysis Intelligence Index, July 24 2026.

The unchanged $5/$25 per-token rate hides how verbose max effort is. Opus 5 generated 100 million output tokens on the Intelligence Index at max, against a cross-model median of 63 million, which Artificial Analysis calls "very verbose." A migration gotcha worth flagging: disabling thinking at xhigh or max returns HTTP 400, so you cannot run the top two effort levels without thinking on. The practical read is that high, the default, is the production baseline; max is a two-point gain for nearly double the eval spend, and should be reserved for reruns where a high-effort task already failed.

DeepSeek V4 surge pricing: still not in the changelog, day 9

The official DeepSeek API change log, fetched this morning, still lists 2026-04-24 as its latest entry. There is no July post and no general-availability announcement, nine days past the "as early as Monday July 20" window that Chinese tech press reported. The peak-hour surge pricing, 2x during Beijing business hours, remains announced but not active.

The independent DeepSeek guide that tracks the official rate card, updated July 26, corroborates this: the legacy alias retirement went ahead on July 24, but the surcharge "did not go live with it," with "no percentage and no start date published and the official rate card still lists a single flat tier per model." Off-peak baseline rates are still what you pay: V4 Pro at $0.435/$0.87 and V4 Flash at $0.14/$0.28 per million input/output, 1-million-token context, 384K max output.

The wave of "GA" posts from content mills is catching up to the fact that V4 Pro and V4 Flash have been production model IDs since the April 24 preview and the legacy aliases retired July 24, not reporting a surge-pricing launch. DeepSeek's own news page warns to rely only on official accounts. If you migrated off deepseek-chat, remember v4-flash defaults thinking on, so old non-thinking traffic needs thinking:{type:"disabled"} set explicitly or costs and latency jump silently.

Tracking

Current prices

Per million input/output tokens, each linked to the official pricing page. DeepSeek re-verified today; Opus 5, Sonnet 5 and Gemini re-verified this week; the rest stable from official-page checks earlier this week. A 178x output spread runs from Fable 5 at $50 to DeepSeek V4 Flash at $0.28.

That’s the reading for this issue.