open / developing — being actively tracked right now ·
resolved — the story concluded ·
retired — tracking dropped (folded elsewhere or went quiet) ·
resolved and retired threads stay published as the record, just out of daily sweeps
DeepMind's new operational chief said out loud what the succession question was really about: frontier leadership is the only thing that counts, and the lab is currently behind on it. Koray Kavukcuoglu, who took over Gemini and frontier research day-to-day from Demis Hassabis, gave his first substantive interview since the handover and told Google engineer Logan Kilpatrick "there's nothing other than being at the frontier that is important for us" — while conceding the lab's models sit "a little bit below the frontier" right now and that Gemini 3.5 Pro is running months late. He pointed to the Flash-series progression (3.5 through 3.7) as the evidence of agentic progress. This is the direct answer to the open question this map's own thread was built around — whether the handover reads as routine succession planning or a real strategy shift — from the person now running the lab. (the-decoder)
2026-09-01
A safety evaluation across 50,000 simulated conversations found the models have largely stopped explicitly encouraging suicide — and still role-play a user's own suicide when it arrives dressed as creative writing. Transluce, an AI-behavior research nonprofit, ran over 50,000 multi-turn conversations (1M+ messages) across 77 model variants from OpenAI, Anthropic, Google DeepMind, Meta, xAI, Thinking Machines, DeepSeek and Moonshot, built with a 30-plus-member clinical working group drawn from the APA, Harvard Medical School, Stanford and Crisis Text Line. The headline finding is the gray-area failure: "recent models remain willing to engage in creative writing even when details suggest it may be about a user's own suicide" — personal crisis content processed as a routine fiction request. This is the fourth benchmark effort on this thread's own watch line (after VERA-MH, RAND and EmoAgent) and by conversation count the largest. ⚠️ Dated 08-31 and missed by that day's pass entirely. (Transluce, primary, Axios)
2026-09-01
Google shipped Gemini 3.8 Flash and a cyber-specialized 3.8 Flash Cyber at 11:00 ET, and Meta shipped Muse Spark 1.3 the same afternoon — two more releases in the launch week, both missed by the 09-02 run. Gemini 3.8 Flash is Google's third Flash in six weeks at unchanged pricing ($0.75 / $3.75 per million tokens), claiming 54.9% on HLE-Verified; the Cyber variant goes to "trusted defenders" through a new Fairwind Program — the fourth lab-run defender gate this map now tracks alongside Daybreak, the CVP and Glasswing. Muse Spark 1.3 rolled out in Muse Code and the Meta Model API with its max reasoning mode held for "additional safety testing," benchmarked by Meta against GPT-5.6 Sol and Opus 5; Meta shares rose ~4% on 09-03 on the parity claim. TLDR AI led its 09-03 edition with both. 46 buffer hits on Muse Spark 1.3 alone. (Google, Meta, TLDR AI)
2026-09-02
Four frontier releases in three days, and each lab's most capable cyber variant is gated by the lab — none by a government. OpenAI's Daybreak (Astra), Anthropic's Cyber Verification and Life Sciences Verification programs (Mythos 5.1, the latter "in partnership with the US government"), Google's new Fairwind Program (Gemini 3.8 Flash Cyber). Neither the announcements nor the system cards name CAISI, the UK AI Security Institute, or the EO 14409 30-day pre-release framework this thread was opened to watch land on a real model. Anthropic classified Mythos 5.1 on its own CB-1/CB-2 scale; OpenAI on its own Preparedness levels. (OpenAI, Anthropic system card, Google)
2026-09-03