August 2, 2026
OpenAI names its next model Astra, which solved 10 open math problems
Subscribe
OpenAI confirmed the Astra name Aug 1 with ten decade-old math proofs formalized in Lean (about $2,000 in compute at Sol rates) but no release date, price, or final model name; DeepSeek posted V4 Flash -0731 open weights under MIT (Artificial Analysis Intelligence Index 50, third among open-weights); MiniMax H3's per-second video pricing landed at $0.13/sec; and AMD shipped a "fully open" 16B MoE under a research-only license.
OpenAI confirms Astra, its next major model, with ten Lean-formalized proofs
OpenAI confirmed the Astra name in public for the first time on August 1, 2026, attaching it to a bundle of ten results in pure mathematics and theoretical computer science that an internal, unreleased build of Astra produced. The company calls Astra its "next major model," and the post is explicit that the arguments came from the model, not humans: "claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system's contribution and the nature of genuine human intellectual work."
The ten results each sat open for at least a decade, and several far longer. The headline is the first explicit construction of a non-sofic group, closing a question Mikhail Gromov opened in 1999 (27 years). Others include a disproof of Connes's rigidity conjecture on von Neumann algebras, a proof of Ehrhart's volume conjecture, three Erdos problems (183, 146, and 180), the first improvement to the general upper bound on high-dimensional sphere-packing density since 1978, a parallel repetition theorem for two-player quantum games, and new lower bounds on the circuit complexity of the permanent.
The pricing beat has a concrete number here. OpenAI says the tokens needed to find all ten solutions would have cost roughly $2,000 at GPT-5.6 Sol API rates. Humans then worked with the same model to turn the raw arguments into a 249-page manuscript, and Astra formalized each proof in Lean 4. The openai/ten-proofs repository pins Lean 4.32.0 and Mathlib, ships Apache 2.0, and includes Comparator configs that recheck the proofs with an external kernel independent of the Lean compiler. A Lean-checked argument is not the same as a peer-reviewed one (no specialist review of the bundle had been published by the Aug 1 cutoff), but it does close the usual failure mode of AI proof claims: a plausible chain that quietly hand-waves a step.
The reason this matters on a releases beat is the availability gate, not the math. OpenAI has not released Astra, has not set a date, has not published a price, and has not even settled whether it ships as GPT-6, GPT-5.7, or a separate class alongside Sol, Terra, and Luna. What it has done is put Astra squarely inside a new federal review apparatus. The Decoder reports that CEO Sam Altman demoed Astra to Washington policymakers the same week, and that Astra is expected to be among the first models tested under a planned US government framework requiring federal clearance before public release. Executive Order 14409, signed June 2, 2026, gave NSA, CISA, and the Treasury a 60-day clock to design a classified benchmarking process for "covered frontier models" plus a voluntary framework for up to 30 days of government pre-release access. That clock ran out on August 1, the same day OpenAI revealed the name. The order stops short of mandatory licensing, but the timing means Astra's public availability now depends on a clearance process that did not exist a week ago.
For buyers: nothing to switch to yet. Astra is an internal research system, and OpenAI frames it as a model built to coordinate multiple agents on long-running problems, not a drop-in replacement for the Sol/Terra/Luna ladder you are calling today. The same week also brought ChatGPT for Academic Researchers, giving 100,000 scientists free access to OpenAI's top-tier models through 2027.

DeepSeek V4 Flash -0731 open weights land on Hugging Face under MIT
The V4 Flash -0731 build that went official on July 31 (same $0.14/$0.28 price, stronger agents) is now downloadable. Tracking registries that re-checked on August 1 report DeepSeek published the 0731 checkpoint to Hugging Face under the MIT License, and moved the three Flash rows to "open weight" status after the re-check. Anyone already pointed at the deepseek-v4-flash API ID is served -0731 automatically with no migration and no new model string; the open weights mean you can also serve the same checkpoint yourself and skip the queue.
The independent benchmark landed too. Artificial Analysis gave the max-reasoning configuration an Intelligence Index score of 50, ranking it third among 101 large open-weight models in its comparison class. That placement matters because DeepSeek charges $0.14 per million cache-miss input tokens and $0.28 per million output tokens (cache hits fall to $0.0028), with a 1 million-token context window and 384,000 max output tokens. At those rates a coding agent that runs long prompts, repeated tool calls, and large reasoning traces inherits none of the premium-proprietary economics.
What did NOT move: V4 Pro and the surge pricing. Only the Flash API was upgraded; V4 Pro is still "coming soon" on the official pricing page (fetched this run, still lists deepseek-v4-flash updated to -0731 and deepseek-v4-pro). DeepSeek has said every billing item will cost double inside two Beijing-time windows (09:00-12:00 and 14:00-18:00, equal to 01:00-04:00 and 06:00-10:00 UTC), implying $0.28 and $0.56, but no effective date has been published. The off-peak baseline is what you pay today. Treat the open weights as the pressure valve: once the checkpoint is public, serving becomes a commodity market across the providers already hosting it.
MiniMax H3 ships per-second video pricing at $0.13/sec, weights still promised
MiniMax's omni-modal H3, which went live July 31, now has a real price card, and the unit is the story: it bills per second of generated video, not per token. The 2K tier is $0.13 per second ($7.80 per minute), with a 768p tier at $0.09 per second in closed beta (contact sales). That $0.13 figure comes from the launch breakdown and third-party trackers; MiniMax's own pay-as-you-go page still listed only Hailuo 2.3 tiers at the Aug 1 cutoff, so treat it as reported until it appears on MiniMax's page. The first five reference images are free, then $0.04 each; reference audio is free; reference video input is billed by input duration at the output tier rate. A 6-second 2K clip with audio costs $0.78.
The fine print for buyers: H3 is pay-as-you-go only. MiniMax's subscription packages explicitly state "MiniMax H3 is not supported yet," so package holders still use the older Hailuo 2.3 tiers. On the Artificial Analysis leaderboards H3 ranks number one in Video Editing, number two in Text-to-Video, and number three in Image-to-Video, which MiniMax says makes it the strongest open-weights video model ahead of the previous leader LTX-2.3 once the weights actually ship.
They have not shipped yet. The announcement says MiniMax plans to open the weights "in the coming days, subject to applicable laws and regulations," but as of August 1 there was no Hugging Face repository or model card, so there is no verified parameter count, VRAM floor, or official benchmark table. The license to watch is the planned MiniMax Community License, which Artificial Analysis reports permits commercial use for organizations under $20M revenue with prominent attribution. That is a materially different deal from the MIT license MiniMax used on its earlier M-series checkpoints, so read the final text when the repo lands rather than assuming MIT.
Open-weights watch: AMD's "fully open" 16B MoE, LG's 750B EXAONE
AMD released Instella-MoE-16B-A3B on August 1, a Mixture-of-Experts model with 16B total and 2.8B active parameters, trained from scratch on Instinct MI300X and MI325X GPUs. The notable part is what "fully open" means here: AMD is publishing weights from every training stage, plus the data mixtures, training configs, and inference code, not just a final checkpoint. Architecture is 27 layers, 64K context (extended from 4K with YaRN), 7.1T pre-training tokens, two shared plus six routed experts from 64, Gated Multi-head Latent Attention, and a FarSkip-Collective connectivity trick AMD says cut pre-training time 12.7% and time-to-first-token up to 39.2% under expert parallelism. AMD reports the base checkpoint averages 76.7, calling it the strongest among "fully open" models.
The license is the catch for the pricing beat. The weights ship under ResearchRAIL, which permits academic and research use only, so this is not a drop-in commercial model; the training codebase is separately MIT licensed and is the more reusable asset. Inference needs roughly 32 GB of weight memory in BF16, so one high-memory accelerator is enough, and AMD ships SGLang inference code. Treat it as a reproducibility and research artifact, not a serving-tier competitor to the MIT-licensed DeepSeek Flash or the Apache-2.0 options below.
LG AI Research also released a 750-billion-parameter EXAONE model as an open-weight asset, framed as South Korea's largest and a sovereign-AI play to cut reliance on US and Chinese stacks. The coverage this run is thin on primary detail: no official license name, no model card link, and no pricing were confirmed from an LG primary source by the Aug 2 cutoff, so treat the parameter count and "open-weight" framing as reported rather than verified, and read LG's own release page before building on it.
Tracking
- Aug 5 double cutover, 3 days out.
grok-voice-latestmoves from Voice 1.0 to Think Fast 2.0 (auto; pin 1.0 to stay, xAI) the same day Claude Opus 4.1 retires and the official models overview now says "migrate to Claude Opus 5" at the same $5/$25. Two breaking changes on the Opus 5 migration: thinking is on by default, and disabling thinking at xhigh/max effort returns HTTP 400. - MAI-Cyber-1-Flash / Project Perception public preview opens tomorrow, Aug 3. Microsoft's first cybersecurity model (137B total, 5B active MoE, 256K context) stays MDASH-only, not a standalone API, priced on consumption in "Security Compute Units" rather than per token. The private preview reportedly found 16 new Windows flaws including four critical RCEs.
- Grok 4.6, Aug 7, 5 days out. Elon Musk's July 28 X post targets Aug 7 for a 1.5T model with improved SFT and RL; Arena.ai confirmed the date and says it will evaluate the model the following week. No pricing, context window, or model card published. Grok 4.7 (2.1T) follows "a few weeks later."
- Sonnet 5, Sep 1, 30 days out. Intro $2/$10 runs through Aug 31, then $3/$15. The new tokenizer produces about 30% more tokens, so the effective Sep 1 rate is closer to $3.90/$19.50, more than Sonnet 4.6's $3/$15.
- Gemini 3.5 Pro, still missing. 75-plus days past the I/O "next month" promise, no
gemini-3.5-proAPI entry, no pricing row. A prediction market puts it at roughly 66% by Aug 16 (treat as speculation, not a Google date). Gemini 3.6 Flash at $1.50/$7.50 is the stopgap; Gemini 4 pre-training has begun. - Ant Ling-3.0-flash is free through Aug 3, with weights promised after. Mistral's frontier MoE is in early access with no name, specs, or price (Large 3 holds at $0.50/$1.50, Apache 2.0). Qwen 3.8-Max-Preview remains credits-only since July 19.
Current prices
Every row links to the vendor's own pricing or models page. OpenAI and DeepSeek rows were re-fetched this run (Aug 2); the rest were verified on their official pages within the last week. Prices are standard short-context API rates in USD per 1M tokens unless noted.
- GPT-5.6 Sol: $5 in / $30 out, 1.05M context, 128K max out, Feb 16 2026 cutoff. developers.openai.com. Sol Fast mode is 2.5x faster at 2x the price ($10/$60), replacing Priority Processing.
- GPT-5.6 Terra: $2 / $12 (cut 20% on Jul 30, from $2.50/$15). developers.openai.com.
- GPT-5.6 Luna: $0.20 / $1.20 (cut 80% on Jul 30, from $1/$6). developers.openai.com.
- Claude Opus 5: $5 / $25, drop-in for the retiring Opus 4.1. platform.claude.com.
- Claude Fable 5: $10 / $50 (cache hits $1, 5m/$12.50, 1h/$20). Top of the rate card. platform.claude.com.
- Claude Sonnet 5: $2 / $10 intro through Aug 31, then $3 / $15 (new tokenizer adds ~30% tokens). platform.claude.com.
- Claude Haiku 4.5: $1 / $5. platform.claude.com.
- Gemini 3.6 Flash: $1.50 / $7.50. ai.google.dev.
- Gemini 3.5 Flash-Lite: $0.30 / $2.50. ai.google.dev.
- Grok 4.5: $2 / $6, 500K context. docs.x.ai.
- Mistral Large 3: $0.50 / $1.50, Apache 2.0. mistral.ai.
- Kimi K3: $3 / $15 (cache $0.30). platform.kimi.ai.
- DeepSeek V4 Flash (-0731): $0.14 / $0.28 (cache hit $0.0028), 1M context, 384K max out. Open weights, MIT. api-docs.deepseek.com.
- DeepSeek V4 Pro: $0.435 / $0.87 off-peak, "coming soon." Surge (2x peak hours) announced, not yet active. api-docs.deepseek.com.

That’s the reading for this issue.
- OpenAI slashes GPT-5.6 Luna 80% to $0.20/$1.20, DeepSeek V4 Flash goes official Jul 31
- Grok Voice Think Fast 2.0 launches at $0.08/min, grok-voice-latest upgrades Aug 5 Jul 30
- Grok 4.6 targets Aug 7 at 1.5T, Grok 4.7 follows at 2.1T Jul 29
- Kimi K3 License is MIT-like, but a $20M Model-as-a-Service clause applies Jul 28
Want the next one?
Every new AI Releases & Pricing issue by email. One tap to unsubscribe.