The Model Stops Being a Chatbot: ChatGPT Work, Claude Cowork, and the Agentic Workflow Race
The most consequential thing about this week's OpenAI launch was not the model — it was the wrapper. Alongside GPT-5.6, OpenAI folded Codex into a revamped ChatGPT desktop "superapp" with a built-in browser and computer control, and shipped ChatGPT Work, an agent that runs multi-step tasks for hours across connected apps like Slack, Google Drive and Microsoft 365. The Codex app becomes the ChatGPT desktop app on Windows and Mac, on every plan including Free. As Forbes framed it, this is a deliberate pivot from a place you chat to a place work gets done — and it landed on 9 July, two days after Anthropic put Claude Cowork on web and mobile in beta, removing the desktop constraint on long-running agentic sessions.
For European enterprises in regulated sectors, the wrapper is where the compliance question actually lives. A chatbot sees one prompt; a workflow agent granted standing access to a company's Slack, its Drive and its Microsoft 365 tenant sees the working material of the whole organisation, and reaches into it repeatedly over hours. That is the data-in-cognition problem moved up a layer — from "what did you type" to "what can the agent read on your behalf." When the reasoning behind those actions runs on a US-hosted model, the CLOUD Act exposure is no longer about a single API call but about a persistent, privileged reader sitting next to regulated data. The agent era makes where the model runs a governance decision, not a procurement footnote.
Three Frontier Launches, Independently Measured — and the Harness Decides the Winner
With GPT-5.6, Grok 4.5 and Claude Sonnet 5 all landing inside two weeks, the useful signal this week came from independent measurement rather than launch-day tables. Artificial Analysis, which pre-evaluated all three OpenAI models, placed GPT-5.6 Sol second on its Intelligence Index — behind Claude Fable 5 — while putting Sol first on the new Coding Agent Index at 80.0, 2.8 points above Fable 5, at roughly one-third the cost and less than half the tokens and time. The lesson is that the leaderboard now depends on which agentic harness you pair a model with; the durable enterprise number is cost-and-tokens per completed task, not a single headline percentage.
Grok 4.5, xAI's first flagship since SpaceX absorbed the company, is the cautionary case. Elon Musk claimed the top spot "on multiple benchmarks"; Artificial Analysis ranked it fourth on the Intelligence Index at 54, behind Fable 5, GPT-5.5 and Opus 4.8. It is genuinely cheap — about $0.31 per Intelligence-Index task — and posted the best agentic tool-use score charted, but its hallucination rate climbed from 25% to 54% even as its knowledge accuracy rose. A model that knows more and is also more confident when wrong is exactly the profile a regulated buyer should test on its own workload before trusting, not adopt on a founder's leaderboard post.
Mistral's Open-Weight Bet: Europe's Sovereign Play Goes to the Weights
Against a week dominated by US and Chinese launches, Mistral is pressing the one lever neither can match for European buyers: open weights under EU jurisdiction. CEO Arthur Mensch has confirmed a new Mixture-of-Experts family he describes as "fat but sparse" entering early access this month, backed by a €4 billion data-centre build-out across France and Sweden — including a €1.2 billion hydropower-backed EcoDataCenter facility in Borlänge — and February's Koyeb acquisition to anchor what Mensch calls "a true AI cloud."
The strategic argument, as Mensch puts it, is "strategic autonomy": the ability to run, inspect and modify a model free of a vendor's terms of service or a foreign government's jurisdiction. On-premise deployment through open weights is the cleanest answer to the CLOUD Act question because the data never leaves the customer's own infrastructure. The timing is not incidental — EU AI Act enforcement powers, including requests for information, model access and recall, activate on 2 August, and the sovereign-deployment argument reads very differently the week regulators gain teeth. It is the same case for open weights that the Chinese labs are making on price; Mistral's version keeps the jurisdiction in Europe.
Quick Hits
- EU Digital Omnibus clears its last hurdle. The Council gave final green light on 29 June (Parliament, 16 June); Official Journal publication is expected imminently, with entry into force three days after. High-risk Annex III obligations slip to 2 December 2027, a new Article 5 prohibition on "nudifier" imagery and CSAM is added, and content-watermarking rules move to 2 December 2026 — but the core enforcement powers still switch on 2 August.
- The US voluntary frontier framework is due 1 August. Sixty days after the 2 June executive order, the NSA must finalise a classified process for designating "covered frontier models" and a multi-agency group must publish the voluntary early-access framework; GPT-5.6's roughly two-week gate for a small set of vetted organisations was the informal dry run.
- Grok 4.5 skipped the EU. xAI cited a regulatory compliance review for the model's absence from the EU at launch, targeting mid-July access — a small, concrete illustration of the 2 August enforcement clock shaping rollout choreography.
- Chinese open weights keep taking share. Chinese open-weight models now sit near the top of neutral rankings and, by recent counts, a majority of tokens routed through OpenRouter, with Z.ai's MIT-licensed GLM-5.2 leading the field — increasingly on Huawei silicon. The price case for open weights is now also a supply-chain-independence case.
- The forward-deployed-engineering arms race compounds. Microsoft's $2.5B Frontier Company (6,000 embedded experts; LSEG, Unilever and Land O'Lakes named as early clients) and AWS's $1B deployment venture keep pushing the industry's answer to stalled adoption toward people inside the building — a model that reads awkwardly next to regulated data when the embedded stack is American.
