Editor's note: A day of proof over promises. Palantir put hard numbers on enterprise AI demand, Alibaba pushed a 2.4-trillion-parameter model into the frontier conversation, and a small evaluation startup got paid for the one thing benchmarks can't measure. The EU AI Act's new enforcement powers sit in the quick hits — the deadline has passed and, absent a named action, it stays there.
Palantir puts real numbers on enterprise AI demand
Palantir reported second-quarter revenue of $1.94 billion, up 93% year over year, with U.S. commercial revenue — the segment most exposed to ordinary corporate AI budgets — accelerating 149%. Adjusted earnings landed at $0.41 a share against a consensus near $0.34, and the company raised full-year guidance across the board: revenue to roughly $8.15 billion, U.S. commercial growth to at least 134%, and adjusted free cash flow to $4.5—4.7 billion. The stock jumped about 12% after hours.
For an industry that has spent 2026 arguing over whether AI spending translates into returns, this is one of the cleaner data points on the “adoption is real” side of the ledger — a deployment-and-integration business, not a model lab, compounding demand from named enterprise customers. The caveat worth holding onto is concentration: the acceleration is overwhelmingly U.S. commercial, and European regulated buyers still weigh where the data and the model actually sit alongside the capability. The number that matters for a CFO isn't the benchmark; it's whether the workload can run under the right jurisdiction.
Alibaba's Qwen3.8-Max joins the frontier — weights to follow
Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter sparse Mixture-of-Experts model with about 95 billion parameters active per token, a one-million-token context window, and multimodal input across text, image, and video. On Alibaba's own launch table the model trades blows with the current frontier — 86.6 on Terminal-Bench 2.1, ahead of Claude Opus 4.8 and behind GPT-5.6 Sol — with leads on a handful of agentic and research benchmarks. As always with a vendor's own numbers, those figures are the lab's, not an independent reading, and are worth treating as a starting point until third-party evals land.
The more consequential detail is distribution. Alibaba is billing this as its first open-weight model in the Max tier, but the weights aren't public yet: today it's API-only, with a pledge that they arrive on Hugging Face and ModelScope next week. If that holds, it would be the largest openly licensed model yet to reach frontier-class scores — the kind of build a European operator could, in principle, run on its own soil rather than rent through a foreign API. Every major open-weight frontier launch this year has come from a Chinese lab, and the gap between “downloadable” and “servable at 2.4T parameters” remains a real-world infrastructure problem.
The evaluation layer gets funded
The startup behind Design Arena raised $7.9 million, led by Index Ventures with Conviction, A*, and Valkyrie participating. The product is deceptively simple — show people pairs of AI-generated designs and images and ask which looks better — but the underlying bet is that subjective quality is now a bottleneck the labs can't automate away. The company says roughly 5.3 million people across more than 190 countries have supplied the human judgments, and frontier labs use that signal to tune outputs that don't just pass a test but actually land with a person.
It's a small round, but it points at where value is migrating as raw capability converges and per-token prices fall: away from the model weights themselves and toward the layers around them — serving, deployment, governance, and the human taste that tells one competent answer from a good one. Benchmarks measure what's verifiable; a growing share of what enterprises actually pay for isn't. That's a useful reminder for anyone building a routing or evaluation stack: the score is a floor, not the product.
Quick Hits
- The EU AI Act's enforcement powers are now live. Since 2 August the AI Office can compel documentation, run its own model evaluations, and order mitigation or market withdrawal, with fines up to €15 million or 3% of global turnover under Article 101 (not the €35M/7% cap, which applies to prohibited practices). The Office says its opening move is “technical compliance dialogues,” and models placed on the market before 2 August 2025 have until 2027 to comply.
- Formula 1 is running agentic AI in production on AWS. F1 built a “Data Accelerator” on Amazon Bedrock AgentCore that cut data-source onboarding from up to eight weeks to minutes — a concrete example of agents doing unglamorous data-operations work rather than demos.
- Zhipu's GLM-5.5 watch continues. The ~1-trillion-parameter open-weight model has been widely expected in August since a June JPMorgan note, but Zhipu still promotes GLM-5.2 and has published no card, endpoint, or date. Rumour until a repo says otherwise.
- The self-hostable field keeps closing on the rented one. The same week the evaluation layer got funded, the open-weight cadence out of China shows no sign of slowing — between Qwen3.8-Max, an expected GLM-5.5, Kimi K3, and DeepSeek V4-Flash, the models an operator can hold keep gaining on the ones it can only rent.
