Jalapeño posts its first benchmarks, and the fine print matters
OpenAI published the first benchmarks for Jalapeño, the inference accelerator it built with Broadcom, on 25 August. Against Nvidia systems it claims 1.5 to 1.9 times more work per watt at peak throughput across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5. On Kimi, the largest model in the set, it reports roughly 1.5x higher peak performance per watt and 3.4x lower end-to-end latency; on highly interactive traffic — where every step's delay compounds — it puts the gain at 2.1 to 4.1 times. The tests ran on InferenceX, a public SemiAnalysis benchmark that measures serving a whole request rather than raw FLOPs (TechCrunch).
Read the provenance before the multipliers. SemiAnalysis, publishing the same day, says every number came from OpenAI: it watched the InferenceX runs in the lab but did not run the full suite itself, and it argues the Blackwell comparison is incomplete because Rubin is the fairer opponent. What SemiAnalysis does assert on its own account is more interesting than the perf-per-watt headline — Jalapeño is a general-purpose inference ASIC, not something tuned narrowly to OpenAI's models. A custom chip that only serves its owner's models is a cost story for one company; one that could serve anyone's is a threat to the price floor everybody rents at. Either way it changes nothing you can buy this year: the chip goes into OpenAI's own fleet from the end of 2026, with production ramping gradually through 2027.
Stolen browser sessions are draining Claude accounts, and MFA doesn't help
Anthropic has been emailing affected customers, BleepingComputer reported on 30 August, to say that commodity infostealer malware on their own machines had lifted active Claude login sessions, which an attacker then used to get into accounts and burn through their usage. "If your usage limits looked like they refilled and then drained while you weren't using Claude, this was likely the cause," the company wrote. It named Vidar, LummaC2, StealC, RedLine and Acreed on Windows, plus Atomic Stealer on a small number of Macs, and stressed the malware has nothing to do with Claude — one affected user traced it to a pirated game. Anthropic is signing affected sessions out, stripping saved payment methods so accounts can't be charged, and refunding what it identifies as unauthorised.
The mechanism is the part worth taking to your security team. A copied session cookie is an already-authenticated session, so the attacker never sees a password prompt or a second factor — MFA is simply not in the path. Anthropic's own warning is blunt about the limits of the fix: "Signing you out of Claude stops the stolen sessions, but it doesn't remove the malware." For a regulated organisation the exposure is not the wasted quota, it is that an AI account often holds conversation history, connected tools and billing details, and it now sits in the same target class as a bank login. Short session lifetimes, device-bound tokens and endpoint hygiene are the controls that matter here; buying inference from a provider inside your jurisdiction does nothing about a compromised laptop.
Everyone is deploying agents. Almost nobody has rebuilt the work around them
Deloitte surveyed 501 senior US business and IT leaders between April and June, all of them at least piloting agentic AI, and published the results on 12 August. Forty-two percent have tested or deployed agents and 43% run them in more than one function — but only 15% have scaled orchestrated, cross-functional multi-agent adoption, and Deloitte notes that many of those deployments sit in low-risk, low-ROI corners. Asked where they are prepared, respondents put business processes dead last: 21% called themselves ready, and just 5% said highly prepared. Vision and strategy was the only area to clear 50%.
The blockers named are not model quality. Seventy-two percent say they lack unified, accessible data; 70% say they cannot trust and govern agents; 67% say integration is too costly and complex. Most organisations are layering agents on top of processes they have not redesigned, because layering pays back sooner, and fewer than a third expect most of their processes to be rebuilt around agents inside two years. Meanwhile 43% expect significant workforce disruption within 12 to 18 months, rising to 72% over two to three years — and half say they are not investing enough to handle it. The governance gap is the one to watch: when 70% of adopters say they cannot govern what they have already deployed, "which model, running where, under whose contract" stops being a procurement footnote.
Quick Hits
- GLM-5.3's weights are out, but the licence is not MIT. Z.ai published the 753B-parameter model on Hugging Face on 28 August under a bespoke glm-5.3 licence — two days after shipping the smaller GLM-5.3-Flash under plain MIT. Open weights and open terms are now separate questions; read the file before you plan a deployment around it.
- The next AI Act prohibition lands on 2 December. Under the Digital Omnibus, AI systems that generate child sexual abuse material, or that depict an identifiable person's intimate parts or sexual activity without consent, are banned from that date unless they carry adequate technical safeguards against producing such material. The ban covers images, video and audio, and reaches deployers as well as providers.
- Claude Code's weekly limits move on 14 September. Anthropic said on 29 August it will make a 25% increase over the original baseline permanent — which, measured against the temporary 50% boost running through 13 September, is a 17% cut from what subscribers have today. Both framings are the company's own.
