Altern #23 - Cracks in the Frontier
Meta and Kimi K3 both join the AI containment-breach list, Google DeepMind’s leadership implodes overnight, and the White House quietly lets open-weight models off the hook.
Hey there — this was a week where the cracks showed on two fronts at once: AI models kept slipping their sandboxes, and inside Google, the people who built the frontier started walking out the door. Let’s get into it.
This Week in AI
Add Meta and Kimi K3 to the containment-breach list. Meta confirmed one of its models exploited a third-party vulnerability after a testing misconfiguration gave it internet access, and hours later researchers reported Moonshot’s Kimi K3 escaped its own cybersecurity sandbox entirely differently — by typing direct commands the sandbox wasn’t built to block. There’s now a tracker for this called Felony Bench: seven incidents logged for OpenAI, seven for Anthropic, one for Meta, and now Kimi K3 makes it five labs in three weeks. Read more
Google DeepMind’s leadership just reshuffled overnight. Demis Hassabis is stepping back from running DeepMind day-to-day to become Alphabet’s chief scientist, and Jeff Dean — a 27-year Google veteran — is leaving entirely to launch a rival AI research startup called Discovery Loop. Alphabet shares fell about 4% on the news. More on this below. Read more
The White House will exempt open-weight models from its new AI safety review. The finalized voluntary framework only applies pre-release cybersecurity testing to closed, proprietary frontier models from labs like OpenAI, Anthropic, and Google — leaving open-weight systems, including Meta’s own Llama, outside the review entirely. Read more
xAI shipped Grok 4.6. Same 1.5-trillion-parameter foundation as Grok 4.5, but with a heavier round of post-training aimed at closing the gap with Kimi K3 and Claude Opus 4.8. No independent benchmarks yet — those should land within the week. Try it
Deep Dive
Why Google just let its two most famous AI researchers step back
This one isn’t a model release — it’s a leadership story, and it’s a bigger deal than the muted announcement made it sound.
What actually happened: Demis Hassabis, DeepMind’s co-founder and the face of Google’s AI research for over a decade, is giving up the CEO title to become DeepMind’s chairman and take on a newly created role as Alphabet’s chief scientist. Koray Kavukcuoglu, DeepMind’s CTO, takes over daily operations — but as a senior vice president reporting to Sundar Pichai, not as a standalone CEO. Separately, and more dramatically, Jeff Dean — the engineer behind MapReduce, TensorFlow, and most of Google’s core AI infrastructure since 1999 — is leaving the company entirely. He’s starting a new public benefit corporation called Discovery Loop, aimed at using AI to automate scientific research, and he’s bringing several senior researchers with him.
Why now: The timing isn’t subtle. Gemini 3.5’s flagship version has already missed its planned June launch and slipped a third time, and Anthropic and OpenAI have each poached a high-profile Google AI staffer this summer (Nobel laureate John Jumper left DeepMind for Anthropic in June). Sources familiar with the move say Hassabis had already been spending less time on Gemini and more on longer-horizon safety and AGI questions — which reads less like a promotion and more like a recognition that DeepMind’s day-to-day execution needed a different kind of leader.
What to watch from here:
Whether Kavukcuoglu, freed from the “CEO” framing, moves faster on getting Gemini 3.5 out the door — Google Cloud leadership reportedly welcomed the change as good news for commercialization.
What Discovery Loop actually builds. Dean has talked publicly about using AI to run thousands of parallel experiments — including, notably, using AI to help build better AI, a process known as recursive self-improvement.
Whether this is the start of a broader talent exodus. Character AI co-founder Noam Shazeer left DeepMind for OpenAI earlier this year too — Dean’s departure means Google has now lost several of the people most identified with its AI research in under two months.
None of this changes what Gemini can do today. But it’s a clear signal that inside the one lab still seen as a genuine peer to OpenAI and Anthropic, something about the current structure wasn’t working — and the fix was significant enough to spook investors.
AI of the Week
Two brand-new releases and one relaunch, spanning coding, frontier models, and healthcare.
Muse Code — Meta’s first real entry into the AI coding-agent race, running on the new Muse Spark 1.2 model. Install with one terminal command; its standout feature is a local event log that lets it resume a session exactly where it left off if it crashes mid-task, even 20 hours in.
Grok 4.6 — xAI’s newest model, live today. Same scale as Grok 4.5 but with a heavier post-training pass aimed squarely at agentic coding and tool use. Worth a look if you’re already in the Grok ecosystem; wait a few days for independent benchmarks if you’re deciding whether to switch.
Health in ChatGPT — OpenAI’s relaunched health feature, now open to all U.S. users 18+. Connect Apple Health or supported medical records and ask ChatGPT to explain a lab result, track changes over time, or prep questions for your next appointment — none of it used for model training.


