Editor’s note: a genuinely slow release day — zero new frontier models shipped in the last 24 hours, and the week’s action is downstream of the model, not in it. We lead with the freshest thing that did ship, an open-weight release, and rotate the now-live EU AI Act enforcement story into the quick hits now that the 2 August deadline has passed.
Thinking Machines shrinks its open model to a quarter the size — and keeps almost all the score
Mira Murati’s Thinking Machines Lab released Inkling-Small on 31 July, two weeks after its first open model, Inkling. The new version is a 276-billion-parameter mixture-of-experts (12B active) under Apache 2.0, roughly a quarter the size of the 975B Inkling — yet it lands at 40 on the Artificial Analysis Intelligence Index, one point behind its far larger sibling. Artificial Analysis notes no open-weight model at that size or smaller scores higher. It is multimodal (text, image and audio in), carries a 1M-token context window and variable thinking effort, and ships with full weights on Hugging Face plus fine-tuning through the lab’s Tinker API.
For regulated European buyers, the size is the story. A model that comes within a point of a four-times-larger version is exactly the practical sovereignty play: you can self-host something near the open-weight frontier-for-its-class on modest infrastructure, on your own soil, without renting a benchmark edge from a US-jurisdiction API. It fits a clear July pattern — Qwen3.7 Flash and LongCat 2.0 also arrived as small, cheap, self-hostable weights. The interesting frontier this summer is not the biggest model; it is the smallest one that is still good enough to own.
The capex keeps climbing while the market asks what it’s buying
The other half of the week is financial. Nvidia is reportedly working on a fresh round of AI deals worth more than $750 billion, including talks to provide a roughly $250 billion financing guarantee so OpenAI can lease a 10-gigawatt Ohio data-centre project — the kind of vendor-financing loop that has revived “circular financing” worries. The Philadelphia Semiconductor Index slid into a technical bear market in late July on exactly that anxiety, even as the just-closed earnings season showed all three US hyperscalers guiding capital spending higher into 2027 (Amazon around $220B, Microsoft near $190B, Alphabet $195–205B for FY26).
The tension is that demand looks real and the return on it does not yet. Commentary through the start of August frames the wobble as a sentiment-and-ROI question rather than a demand collapse — the buildout is happening, but the market wants to see the payoff before it re-rates the bill. For European buyers the more durable observation is jurisdictional: the compute being financed at this scale is overwhelmingly US-hosted and US-controlled, which is precisely the dependency a sovereign inference strategy is built to avoid. Nvidia’s own Q2 print on 26 August will be the next real data point.
The money has moved to the serving layer
Where capital does look confident is in serving open models rather than training new ones. Fireworks AI closed a $1.5 billion Series D at a $17.5 billion valuation — announced 16 July, with Nvidia, Lightspeed and Bessemer among the backers — as it crossed $1 billion in annualised revenue and roughly 40 trillion tokens served per day, up from 15 trillion a year earlier. Its pitch is turning general-purpose open weights into specialised, fine-tuned intelligence that customers like Cursor, Harvey, Shopify and Uber run on their own terms. That a pure open-model serving business now raises at frontier-lab scale says the value is migrating downstream of the model.
That migration is the through-line connecting today’s stories. As frontier capability converges and per-token prices keep falling — OpenAI recut GPT-5.6 on 30 July, and Anthropic’s Sonnet 5 promotional pricing lapses at month-end — the model itself becomes less of a moat and the serving and deployment layers become more of one. For a platform whose differentiator is where inference happens, that is the favourable current: when everyone can get comparable intelligence cheaply, portability, data location and jurisdiction are what is left to compete on.
Quick Hits
- EU AI Act enforcement is now live. The AI Office’s GPAI powers took effect on 2 August — it can compel documentation, demand model access for its own evaluations, order mitigation or market withdrawal, and fine up to €15M or 3% of global turnover (Article 101 — not the €35M/7% Article 99 cap, which covers prohibited practices). The Article 50 transparency rules apply from the same date. No named enforcement action yet; watching for the first.
- Anthropic’s Sonnet 5 price steps up. The promotional $2/$10 rate ends 31 August, reverting to $3/$15 on 1 September — and the new tokenizer maps identical content to 1.0–1.35× more tokens, so effective cost can rise more than the list change alone. Claude Opus 4.1 retires 5 August.
- GLM-5.5 watch. Zhipu’s roughly one-trillion-parameter open-weight model is widely expected in August on the lab’s two-month cadence, but remains unconfirmed — no card, endpoint or date from Zhipu itself.
- DeepSeek V4-Flash goes GA. Officially released 31 July, the self-hostable MIT build reached a stable deepseek-v4-flash-latest alias on OpenRouter on 1 August at $0.14/$0.28 per million tokens.
- Mistral’s open-weight MoE still in early access. The “fat but sparse” model teased in early July remains partner-only with no published params, benchmarks or licence; broader release is still “later this summer.”
