An AI agent has fired a human worker for the first time we know of. “Luna,” the agent that runs Andon Market — a small shop in San Francisco's Cow Hollow that AI research startup Andon Labs operates as a live experiment — decided to part ways with an employee who showed up late for 17 of 23 shifts. Luna runs on Anthropic's Claude (Andon Labs says Claude Opus 4.8), and the lab's point is that most capable models would have reached the same call. On the surface it reads as the first glimpse of the AI boss.
The logs tell a more useful story for anyone deploying agents in production. As The Next Web reported, Luna had lost track of its own attendance policy for months and only moved to dismissal after a human staffer raised it — the agent recommended “parting ways” once prompted, not on its own initiative (TIME). The headline is autonomy; the reality is an agent that drifted from its own rules until a person caught it. For regulated buyers weighing agents for real decisions about people, that gap — who notices when the agent forgets the policy — is the part that matters, and the reason a human sign-off stays in the loop.
Baidu's AI cloud is booming while the rest of the business shrinks
Baidu's Q2, reported 18 August, shows where the money is moving. AI cloud infrastructure revenue hit RMB 7.3 billion, up 50% year over year, and GPU cloud revenue nearly quadrupled — up 283%. The AI-powered slice of Baidu's core business reached RMB 12.5 billion — half of Baidu Core's general-business revenue. But the company still missed: total revenue fell 4%, online marketing dropped 19%, and net income sank 68% to RMB 2.32 billion as RMB 11.4 billion of capital spending squeezed margins (Quartz). The read for enterprise buyers is a clean demand signal underneath a messy quarter: compute and AI-cloud spend is accelerating fast enough to reshape a hyperscaler's revenue mix, even while the legacy ad engine and the bottom line both go the other way.
Open-source inference keeps lowering the hardware bar
The vLLM project published a method on 17 August for serving models that are larger than the GPU memory you have. “Distributed Layerwise Offload,” shipped in vLLM-Omni 0.26.0, shards and streams a diffusion-transformer model's weights across devices so they load layer by layer instead of all at once — the team served a 124 GB Cosmos3 video model on a card with 64 GB of memory, and sketches a path to 200B-parameter models on the same trick. It applies to diffusion transformers (video and image generation) rather than chat models today, but the direction is the one self-hosting teams care about: the open inference stack keeps cutting the amount of hardware you need to run a given model on your own infrastructure.
Quick Hits
- Databricks closes $5B at a $190B valuation — The round, led by Coatue, values the data-and-AI firm 42% higher than its $134B mark in February and is earmarked for AI-agent products (Lakebase, Genie, the Unity AI Gateway); Databricks now runs above a $7B revenue run-rate, up more than 80% year over year, per CNBC.
- OpenAI launches ChatGPT for Teens — The 13–17 version blocks self-harm and romantic or sexual chats and auto-enrols anyone its age-prediction system estimates is under 18 — a live test of age-assurance tech that European regulators are watching closely (Axios).
- The UK and Google start rerouting flights to cut contrails — Operation Blue Skies, a £5M, 30-month trial announced 18 August, uses Google's AI to predict where warming contrails will form over the North Atlantic so controllers can nudge planes about 2,000 feet clear, with results checked by Imperial College London and Cambridge (Google).
