DeepSeek shipped a re-post-trained build of its budget model that now outscores its own flagship on agent tasks, the AI industry's biggest names keep pouring money into deployment services rather than new models, and the EU AI Act's enforcement teeth come out on Monday. Today's throughline: capable models keep getting cheaper — and where you can actually run them keeps mattering more.
DeepSeek's cheap model overtakes its own flagship — open weights and all
DeepSeek released the public beta of V4-Flash-0731 on Thursday, a re-post-trained version of its 284B-parameter budget model that the company says now beats its own larger V4-Pro-Preview on all nine of its agent benchmarks — at $0.14 per million input tokens and $0.28 output. The headline gain is on Terminal-Bench 2.1, a long-horizon agent test, where DeepSeek's own table reports the score jumping from 61.8 on the preview build to 82.7. The build ships with an OpenAI-compatible encoding format and scores 50 on the Artificial Analysis Intelligence Index, competitive with far more expensive closed models. When one of the cheapest models on the market can hold its own on agent work, the “best model wins” framing keeps giving way to “best fit wins.”
The part enterprise buyers should note is what came with it. Unlike much of this summer's “announced but not downloadable” pattern, DeepSeek published the 0731 build itself to Hugging Face under an MIT license — a dedicated deepseek-ai/DeepSeek-V4-Flash-0731 repository, with community quantized (GGUF) versions already circulating for local use. In other words, the version that tops the benchmarks is the version you can self-host, not just the one behind DeepSeek's API. For a regulated European team, that is the whole game: a benchmark-leading, rock-bottom-priced model whose weights can run on EU soil under your own governance is exactly the artifact sovereign deployment is built to serve — the caveat being that at 284B parameters, “self-host” still means a serious GPU node.
Enterprise AI's money moves from models to deployment
The clearest enterprise signal this summer is not a model — it is where the money is going. The industry's biggest names are pouring capital into enterprise deployment ventures, betting that the next competitive battle is implementation, not raw capability. Microsoft's $2.5B Frontier Company, announced 2 July with 6,000 industry and engineering experts and early clients including the London Stock Exchange Group and Unilever, is the flagship example; AWS stood up a $1B Forward Deployed Engineering organisation two days earlier. OpenAI and Anthropic have launched comparable delivery units of their own. The premise across all of them: capable models are now abundant, and the bottleneck is getting them into production against real workflows and governance.
Europe got its own version of this last week. Cognizant launched a dedicated EMEA AI Unit on 28 July, built explicitly around the IDC finding that roughly 88% of enterprise AI-agent pilots never reach broad production. Its “Frontier Deployed Engineering” offering is tiered from strategy and governance through to end-to-end multi-agent delivery, and Cognizant frames it as independent of any single model, platform or cloud. That platform-neutral, governance-first posture is exactly the ground sovereign infrastructure competes on — the hard part for a European enterprise was never finding a capable model, it was deploying one under a jurisdiction and control model its compliance team can sign off.
EU AI Act GPAI enforcement powers go live Monday
Editor's note: we lead with models and enterprise today, but Monday's deadline is a genuine hard date, so it earns one of the three main slots.
On Monday, 2 August 2026, the European Commission's supervision and enforcement powers over general-purpose AI (GPAI) model providers take effect — one year to the day after GPAI obligations first applied. The obligations themselves are not new; what changes is that the AI Office can now act on them. From Monday it can compel technical documentation, demand access to models to run its own evaluations, order risk-mitigation measures, and restrict or withdraw a model from the EU market. Fines run up to €15 million or 3% of global annual turnover, whichever is higher — the Article 101 cap for GPAI providers, not the higher €35M/7% figure (Article 99) that applies to prohibited-practice violations.
May's provisional “Digital Omnibus” agreement, which pushes back several high-risk AI obligations, deliberately leaves the GPAI track untouched, so Monday's date stands. The practical shift is jurisdictional: the question of who can be compelled to hand over documentation or grant model access, and under which legal regime they sit, moves from a compliance abstraction to an operational one. Models placed on the EU market before 2 August 2025 have until August 2027 to conform; anything newer is already exposed.
Quick Hits
- GLM-5.5 reportedly lines up for August — A JPMorgan research note relayed by Reuters points to a Zhipu AI trillion-parameter open-weight flagship this month, though Zhipu's official channels still promote GLM-5.2 and no model card or endpoint has appeared. It would compete chiefly with Kimi K3, DeepSeek V4 and Qwen — part of Zhipu's stated push toward frontier-tier open weights (“Open Fable”) by year-end.
- Microsoft and Mistral expand their partnership — The 21 July agreement pitches “frontier AI enterprises and regulated industries can control” — a control-and-jurisdiction framing aimed squarely at the same European buyers weighing sovereign options.
- The benchmark ceiling keeps rising — Every frontier model now clears ~88% on MMLU, rendering it saturated; harder tests like Humanity's Last Exam still hold the best models near ~35%. DeepSeek's 50 on the Artificial Analysis Index at $0.14/M input is the more telling number — capability is converging while price collapses.
