Kimi K3's weights go public today, and the real test is the licence, not the parameter count
Moonshot AI set its countdown to end today: the full weights of Kimi K3 — 2.8 trillion parameters, a mixture-of-experts design that fires only 16 of 896 experts per token for roughly 50 billion active parameters, a one-million-token context window and native vision — are due to publish on Hugging Face on 27 July. If they land on schedule it is the largest open-weight model anyone has shipped, and it arrives with genuine standing: third-party trackers place K3 as the top open-weight model on the Artificial Analysis Intelligence Index — fourth overall, behind three closed American systems (Anthropic's Opus 5 and Fable 5 and OpenAI's GPT-5.6 Sol) — with a first-place result in one frontend-coding arena. As of late last week independent trackers still classified it as hosted-only — the API has answered calls since 16 July — so the artifact itself is the thing to watch today.
For anyone deciding what to build on, two numbers matter more than the headline. The first is 1.4 terabytes — the model's MXFP4 weight footprint, which has to sit resident in fast memory before a single token of context loads, putting real self-hosting out of reach of anything short of a full node of Blackwell or MI400-class silicon. The second is the licence, which Moonshot has still not published and says will land with the weights. “Open weights” has stretched from a permissive MIT-style grant to research-only terms with commercial carve-outs, and that difference is the entire question for a regulated buyer weighing sovereign control against a rented API. Downloadable is not the same as runnable, and neither is the same as commercially usable — the sensible posture this week is to prototype and wait for the text.
The memory wall gets its invoice: Korea's chipmakers lock in multi-year supply
That 1.4-terabyte figure is a preview of the constraint now reshaping the whole supply chain. Nvidia used its San Francisco event to unveil a partnership with SK Group worth more than $500 billion, anchored on high-bandwidth memory supply and large-scale data centres slated to come online in 2027. Seoul's Economic Daily reports the wider round of deals — SK hynix and Samsung converting memory contracts with Broadcom, Nvidia and Microsoft into five-year commitments — totals roughly $950 billion, with SK hynix taking about $750 billion of it. The precise aggregate will settle as the contracts are confirmed, but the direction is unambiguous.
The signal for European buyers is that the binding constraint on frontier AI has moved from raw compute to memory capacity and bandwidth — the ability to keep enormous models warm and fed. Locking multi-year HBM supply is now a strategic act on the scale of a national infrastructure programme, and it concentrates that capacity in a short list of hyperscalers and their memory partners. A model can be “open” on paper and still be exercisable in practice only by whoever can afford to hold it in memory at scale, which is precisely why sovereign inference is an infrastructure question before it is a licensing one.
Anthropic's Opus 5 tops the intelligence index at half the price — and the benchmark now measures cost
While the largest open model prepares to ship, the closed frontier answered on economics rather than raw capability. Anthropic's Claude Opus 5, released 24 July, leads the Artificial Analysis Intelligence Index at 60.7% — ahead of Claude Fable 5 (59.9%) and GPT-5.6 Sol (58.9%) — while listing at $5 per million input tokens, half of Fable 5's price. On Artificial Analysis's agentic knowledge-work benchmark AA-Briefcase, it leads by nearly 150 Elo at max effort while cutting cost per task by around a fifth; at the “high” tier a task runs about $10.41 versus Fable 5's $22.30.
The more durable story is the yardstick, not the model. The index's v4.1 revision now reports cost, time and tokens per task alongside the accuracy score, turning the leaderboard into something closer to a unit-economics sheet — and the model that wins on that sheet won on price, not on a new capability ceiling. For a platform choosing which model to route a given request to, that reframing is the point: the frontier contest is increasingly about the cost of a completed task, and the right answer varies request by request rather than settling on a single “best” model.
Quick Hits
- EU AI Act enforcement powers go live in six days. — From 2 August the Commission's AI Office can request documentation, run model evaluations, order mitigations, and fine GPAI providers up to €15M or 3% of global turnover under Article 101 — the end of the one-year grace period, not a new obligation.
- The open-weight calculus cuts both ways. — As China ships the largest open model yet, Beijing is reportedly weighing restrictions on overseas access to its most advanced systems from Alibaba, ByteDance and Z.ai — first reported by Reuters this month and revisited by The Wire China — a reminder that “open” can be revised by the jurisdiction that hosts it.
- Mistral teases a “fat but sparse” open-weight MoE. — Arthur Mensch says the new mixture-of-experts family is in early access for research, government and industry partners, with a broader release later this summer; no parameters, benchmarks or licence yet.
- Gemini 3.5 Pro slips a third time. — Google's flagship remains in partner testing after coding performance fell short of internal targets, with no new public timeline; pre-training on Gemini 4 has reportedly begun.
- DeepSeek completes its V4 cutover. — Legacy deepseek-chat and deepseek-reasoner were retired last week, routing all traffic to the open-weight V4-Flash with a one-million-token default context.
