Kimi K3's open weights arrive Monday, and the story is the hardware bill
Moonshot AI's Kimi K3 — a 2.8-trillion-parameter mixture-of-experts model that has been live via API since 16 July — is scheduled to release its full open weights on 27 July, expected under a modified-MIT-style license whose terms Moonshot has not yet published, making it the largest openly downloadable model in the world. On the Artificial Analysis Intelligence Index it sits at the top of the open-weight field and within striking distance of the closed frontier, with native multimodal input and a one-million-token context window. For anyone tracking whether Europe can run frontier intelligence on its own soil, this is the most consequential drop of the summer so far.
The catch is the 1.4-terabyte weight file. Even in MXFP4 four-bit precision, K3 needs roughly 1.4 TB of fast memory just for the weights, and Moonshot's own guidance points to supernode configurations of 64 or more accelerators for production serving. A modest four-GPU H100 box (320 GB) can technically load a trimmed version, but only with the context window cut to a fraction of its rated length and generation slowed to a crawl. “Open weights” and “self-hostable” are not the same sentence at this scale — the license is free, but the standing infrastructure to serve K3 well is the province of clusters, not laptops. That gap is precisely where a sovereign inference layer earns its keep: the value is no longer just access to the model, it's access to the machines that can run it under European control.
Artificial Analysis rebuilds its benchmark around agents — and retires a saturated test
On 24 July, Artificial Analysis shipped v4.1 of its Intelligence Index, the synthesis score much of the industry uses as shorthand for “how smart is this model.” The headline change is a decisive tilt toward agentic work: Terminal-Bench was upgraded to 2.1, the telecom agent test τ²-Bench became the harder τ³-Bench Banking, and GDPval-AA moved to a v2 revision — all built around messier, multi-step tasks that better resemble how models are actually deployed. IFBench, an instruction-following test, was dropped outright because frontier models had saturated it to the point of uselessness as a discriminator.
Just as telling is what v4.1 adds alongside the score: three per-task operating metrics — cost per task, time per task, and tokens per task — reported for every model. That reframes the leaderboard from a pure capability ranking into something closer to a unit-economics sheet, which is the right lens for anyone actually paying inference bills. At the top, Claude Opus 5 leads at 60.7%, ahead of Claude Fable 5 (59.9%) and GPT-5.6 Sol (58.9%) across 167 models tested. The broader signal for buyers: the questions that separate models are shifting from “can it answer” to “can it finish the job, at what cost, in how many tokens.”
The enterprise money is moving to implementation, not models
The most striking enterprise AI story of the month isn't a model at all — it's the roughly $8 billion that Microsoft, OpenAI and Anthropic have collectively committed to enterprise deployment ventures whose whole purpose is to get AI actually working inside customer organizations. The flagship is Ode with Anthropic, a $1.5 billion services firm publicly launched 15 July by Anthropic, Blackstone and Hellman & Friedman — each contributing roughly $300 million, with Goldman Sachs anchoring another ~$150 million. It fields around 100 engineers, many of them former startup founders, who embed inside enterprises and rebuild core processes around AI, led by the Fractional AI founders Anthropic acquired in May. Microsoft's parallel $2.5 billion Frontier unit and a reported $1 billion from Amazon round out the pattern.
The reason all this capital is chasing services rather than raw capability is blunt: the models are good enough, and the bottleneck has moved. PYMNTS' enterprise benchmarking found that 71% of executives at billion-dollar-plus companies name organizational readiness — not the technology — as the primary barrier to AI performance; only 11% blame the tech itself. For a European platform, that reframes the sale. If deployment, governance and data control are where value now accrues, then where the inference runs and under whose jurisdiction it sits stops being a footnote and becomes part of the readiness problem itself.
Quick Hits
- EU AI Act GPAI enforcement goes live in a week. From 2 August, the EU AI Office can demand documentation, run evaluations, order mitigations and fine general-purpose model providers up to €15 million or 3% of global turnover under Article 101. The obligations bind models placed on the EU market after 2 August 2025; older models get until 2027.
- Mistral teases a “fat but sparse” open-weight model. CEO Arthur Mensch has put a new mixture-of-experts family into early access without disclosing parameter count, benchmarks or license, with a broader release expected later this summer — Europe's best shot at an open frontier entrant.
- Mistral's European build-out targets 200 MW by end-2027. Backed by its expanded Microsoft partnership (21 July) and €722 million in debt financing for a first large-scale data centre near Paris, Mistral is trying to put sovereign compute behind the sovereign-model pitch.
- WAIC establishes a new intergovernmental AI body. July's World Artificial Intelligence Conference in Shanghai (17–20 July) opened with the launch of WAICO, a China-headquartered organisation for global AI cooperation — a governance counterweight worth watching as the EU's own rules take effect.
