Editor's note: today's lead is a regulation story — an exception we make when a single regulatory event genuinely dominates the cycle. China's anthropomorphic-AI rules take effect today, and with them a whole product category goes dark.
The Day the Agents Went Dark
China's Interim Measures for AI Anthropomorphic Interactive Services become enforceable today, and the two biggest consumer AI platforms in the country have complied by deletion. ByteDance's Doubao takes its custom-agent features offline today, leaving users read-only access to their agent configurations and chat histories until 15 October, after which the data will no longer be accessible or recoverable. Alibaba's Qwen switched off human-like and user-created agents on 10 July and retires its broader agent functions today, with configurations and conversation histories permanently deleted and no migration path offered. The rules — co-issued in April by the Cyberspace Administration and four other agencies — ban companion services for minors, mandate non-human labelling, and require anti-addiction systems and instant-exit mechanisms that are architecturally incompatible with persistent-memory agents; workplace assistants and customer-service bots are explicitly excluded.
Three regulatory models are now running in parallel — EU transparency obligations, US access-gating, Chinese product limits — and today the product-limit model produced its first mass casualty: not a fine, but the overnight deletion of a product category, including months of user-accumulated agent state. That is the detail European enterprises should sit with. An agent is not a stateless tool; it is configuration plus accumulated memory, and when that state lives in a vendor's cloud under a regulator's jurisdiction, both the agent and its memory can be switched off with no export path. Persistence and exportability of agent state — where the agent's memory physically lives, and under whose law — just became a procurement question, not a technical footnote.
Production Has Already Left the Frontier
While the industry watched Washington decide who gets access to frontier models, TechCrunch's read of the deployment data makes the counterpoint: the real race may no longer be at the frontier. Chinese open-weight models took 41% of Hugging Face downloads this spring, surpassing US models for the first time, and the six most-used models on OpenRouter are all open models from Chinese firms — Tencent, Xiaomi, DeepSeek, MiniMax and Z.ai. Half of Fortune 500 companies now use Hugging Face to deploy private or open models, with a new repository created on the platform every seven seconds. CEO Clem Delangue's framing is blunt: frontier models are for experimenting and some high-value tasks, while most production workloads run on private or open models — in his words, companies are "done renting their AI".
This is the quiet inversion under the loud frontier news: access to the newest closed model is contested, gated and politically fraught, and enterprises are responding by consolidating production on weights they hold. Note the asymmetry with today's lead — Chinese consumer agents can be deleted by decree, but Chinese open weights, once downloaded, answer to whoever runs them. For a European buyer the lesson cuts both ways: open weights are only a sovereignty win if you self-host them on infrastructure you control, and the 41% download figure suggests that is precisely how they are being consumed. The deployment layer — which weights, whose infrastructure, what routing — is where both the value and the control now concentrate.
OpenAI Puts Its Flagship on a Wafer
OpenAI is deploying GPT-5.6 Sol on Cerebras wafer-scale hardware at up to 750 tokens per second, with access starting for select customers this month — roughly ten times the 40–120 tokens per second a frontier-class model typically streams from a GPU cluster. The deployment sits under a disclosed $20 billion multi-year inference contract and gives the flagship a serving lane distinct from OpenAI's Broadcom "Jalapeño" custom-silicon path. Outside analysts estimate Sol is served across 70–100 wafers at roughly one model layer per wafer — implying a model on the order of 3T total parameters with ~150B active — though OpenAI has not confirmed the architecture.
Speed is becoming a competitive axis in its own right, and the reason is agents: agentic workflows are serial, so latency compounds with every step in the chain, and a 10x throughput gain is the difference between a coffee-break agent run and an interactive one. The infrastructure read matters more than the benchmark: frontier inference is no longer synonymous with Nvidia GPU clusters. Wafer-scale, custom ASICs and on-premises racks are splitting model serving into distinct lanes with different latency, cost and jurisdiction profiles — which means the procurement question is shifting from "which model" to "which serving lane, on whose silicon, in whose datacenter". That is the same question the sovereignty argument has been asking all along, arriving now through the performance door.
Quick Hits
- The EU's Digital Omnibus is still waiting on the Official Journal. Publication is expected mid-to-late July, with entry into force three days later. New details from the final text: the Article 50(2) AI-content-marking obligation gets a four-month grace period to 2 December 2026 for systems already on the market before 2 August, and a new prohibition on AI generation of non-consensual intimate imagery and CSAM lands in December. The 2 August GPAI penalty clock is unmoved.
- Anthropic is talking to Samsung about its first custom chip. Early-stage discussions cover a 2-nanometer inference chip using Samsung's manufacturing and packaging; no design work has begun. Samsung participated in Anthropic's $65B Series H at a ~$965B valuation — the same serving-lane diversification pattern as OpenAI's Cerebras and Broadcom moves.
- The UN's scientific panel says the quiet part in writing. The Independent International Scientific Panel on AI launched its preliminary report out of the Geneva dialogue: with growing evidence of deceptive model behaviour, science currently cannot guarantee that increasingly capable AI will not cause catastrophic harm. The measured institutional language is the story — this is the UN's own evidence panel, not an advocacy group.
- Open-weight watch: Europe's July deliveries are due. Mistral's confirmed open-weight MoE remains in gated early access with no public checkpoint, while OpenEuroLLM's first models, dataset and eval code are due 31 July — two weeks out. Whether Europe's open-model summer ships on schedule is the next test of the supply thesis behind it.
