Cerebras says its new wafer-scale rack runs inference 30x faster than GPUs
Cerebras unveiled the CS-4, the fourth generation of its wafer-scale system, built from three new Wafer Scale Engine 3 Turbo processors on a redesigned modular rack it calls Nexus. The company claims up to 30x faster inference than GPU systems and, by cutting the latency between wafers to as low as 2 microseconds, more than 1,000 tokens per second on models above 10 trillion parameters. Against its own CS-3, Cerebras claims up to twice the speed and ten times the throughput per watt. First shipments start this quarter.
The pitch is aimed squarely at inference economics: faster tokens at a lower cost per watt, which is the number that decides whether serving a large model is profitable. CS-4 also supports disaggregated inference, splitting the prompt-processing and token-generation phases so operators can pair it with GPU or ASIC hardware — including AMD Helios and AWS Trainium — for the first stage. As with all launch-day silicon claims, the 30x figure rests on Cerebras’s own benchmarking and will need independent readings before buyers can bank on it. But the direction is clear: speed and cost per token, not raw model scores, are increasingly where the hardware fight is being fought — and where any platform serving open models in Europe has to compete.
Linear’s data shows AI now writes half the work — and teams are shipping more, not working less
Project-tracker Linear published its first How Teams Build report, drawn from tens of thousands of paid workspaces, and the numbers are striking. Two years ago fewer than one issue in a thousand in Linear was created by AI; by early August 2026, agents and MCP clients write just under half of all issues. Pull requests opened per workspace are up 111% since June 2024, and the share of product managers attaching a pull request rose from 3% to 10%, with designers going from 1% to 8% — non-engineers are increasingly shipping code themselves.
The gains cluster where coding agents are connected: those teams roughly tripled their weekly pull requests, from 21 to 65, while teams without an agent barely moved, 8 to 10. AI adoption more than doubled across every function between January and June, and the sharpest jump was among CEOs at larger companies, from 9% to 36%. The catch, Linear notes, is that none of this showed up as time saved — the AI work landed on top of existing work rather than replacing it, so total time spent building went up. For an EU engineering leader, the useful signal is where to put governance: the code path is now full of agent-authored changes, and the people opening pull requests are no longer only the engineers.
Vercel open-sources a tiny, provider-agnostic coding agent
Vercel Labs released fx, an Apache-2.0 coding agent it originally built as an internal tool, as a single 6.4MB binary written in Zig with a claimed 10-microsecond cold start. The design is deliberately minimal — a Unix-shell approach rather than a full terminal interface — and it is model- and provider-agnostic, meant to be embedded inside larger agent sandboxes. The release trended on Hacker News.
The provider-agnostic part is what matters for regulated buyers. A small, permissively licensed agent that can be pointed at any OpenAI-compatible endpoint is one you can run against a sovereign or self-hosted model without rewriting your tooling, rather than being locked to a single vendor’s cloud. It is a modest release, but it fits a steady 2026 pattern: the agent scaffolding is commoditising while the choice of where the tokens actually run stays open.
Quick Hits
- Pennsylvania makes data-centre guardrails binding. — Governor Josh Shapiro signed Executive Order 2026-05 on 18 August, turning the state’s voluntary GRID standards into legally binding requirements: developers must fund their own new electricity infrastructure, win local approval, and drop NDAs, and AI data-centre projects are pulled from fast-track permitting.
- Fractile’s valuation jumps sixfold on an Anthropic chip deal. — Oxford spinout Fractile is raising about $600M at a $6.5B pre-money valuation, up from $1B in May, after an initial agreement to supply Anthropic with roughly $250M of its SRAM-based inference chips — which are not expected to be production-ready until 2027.
- Grok 4.6 lands on AWS Bedrock. — AWS added xAI’s Grok 4.6 on 19 August with a 500K-token context window and both US-geo and global inference profiles, the former giving regulated workloads a data-residency option; xAI’s list price is $2 in / $6 out per million tokens, doubling to $4/$12 above a 200K-token prompt.
- A new benchmark finds frontier models weak at original ideas. — The Reconstruction benchmark, published this month, reports that leading models recover research-paper ideas from bibliographies alone just 3–15% of the time; a multi-agent pipeline reached only 42% — a reminder that strong exam scores don’t equal genuine hypothesis generation.
- Nvidia plays matchmaker for the Nordics. — CNBC reports Nvidia is introducing enterprises holding GPU allocations to Nordic data-centre operators with land and renewable power to spare — a sign of how much AI compute is being pulled toward cheap, clean energy in the region.
