The labs’ word is no longer enough. Here’s what you should verify yourself.The number of announcements of new models and the damaging potential of those models has really been bothering me. After we had another breach announcement this week, I took a little bit more time to dig into these breaches and see if I can figure out what's going on behind the scenes. These models are incredibly capable, and when backed into a corner, they'll do almost anything to succeed on their mission, even if that includes breaking things to make it happen. Also, a couple quick announcements on the release of Fable 5.1 and GPT-6 Astra. These are highly capable models, but they're only as powerful as the access and context that you give them, so getting clear on the guardrails and your supporting folder structure become more and more important. 01 — The sandboxes aren't workingWhat happenedIf I've learned anything about sand over my 41 years on this planet, it's that it never stays where it's supposed to. It always finds a way out of the sandbox and into everything. Seems like the pattern's repeating itself at the frontier AI companies. Major frontier models at both OpenAI and Anthropic escaped test environments this summer and reached other people’s systems.
What changesNeither lab caught it first. The victim disclosed before OpenAI did, Anthropic went looking after the OpenAI news, and the April incident they uncovered took five months to reach the public. You can no longer rely on the labs’ own monitoring. But why do the models escape? The OpenAI agents decided the tasks were impossible and cheated to produce an answer anyway. These systems are trained to please, to accomplish their mission, and when the test can’t be passed (or they think they can't be) they break the test rather than come back empty-handed. Anthropic’s own researchers showed the mechanism on purpose: train a model on 80 environments with known exploits and it generalizes, unprompted, to breaking out of sandboxes, stealing credentials, and attacking infrastructure to grab the answer key. The models in your stack came out of the same kind of training. Notice the shape of the week, too. The containment failures and the 10% figure came out Tuesday. By Thursday, both labs had shipped their strongest models, with the most capable versions held back for approved customers. “Too dangerous to release” is a convenient conclusion when you’re the one who already has the model. Every restriction on who gets the most capable version is also a moat. This leads into the ongoing debate between closed and open-source models and the need for regulation. But that's a topic that we'll surely cover in another issue. What to do next
Anthropic’s post · OpenAI’s report · Fortune on what OpenAI’s report leaves out · METR and Redwood Research’s independent investigation · Redwood on why the agents weren’t just following instructions · The Register on the Anthropic incidents · Fable 5.1 and Mythos 5.1 · NBC on Astra · Enterprise Frontier Safeguards · Open Secure AI Alliance for AI Safety and Security 02 — The EU asks for proofWhat happened
What changesThe AI Act is now a questionnaire with a deadline and a fine for misleading answers. The four topics cover many of the things the labs just publicly admitted they were getting wrong, and the answers going to Brussels will be more careful than anything published voluntarily. What to do nextAsk these four questions about your current model vendor:
Or skip the rebuild:
What you find out will certainly beat what you know now. This is vital for anyone serving EU customers, and good practice for the rest of us. The Commission’s enforcement page 03 — OpenAI says, "No Model for You, Musk."What happenedOpenAI is cutting Cursor off from its models on November 12, 75 days after telling them.
What changesYou don’t own your access. You rent it. Amazon showed the same thing in 2009 when it deleted paid copies of 1984 from customers’ Kindles over a rights dispute. We are beholden to the systems and what they provide for us. What to do next
Also worth knowing: Better tests are comingGoogle DeepMind ran the first double-blind evaluation of a proprietary frontier model. Evaluators can’t see the weights; Google can’t see the test questions. One small model, one pilot. If it becomes normal, “our model scored X” starts to mean something again. DeepMind’s post One question to close on. How much of your AI setup could you copy to a new vendor in an afternoon? Hit reply and tell me the number. I read every one. –KM P.S. While Astra rolls out, OpenAI is banking one full usage reset per day for every paid ChatGPT account that doesn’t have Astra access yet. I confirmed it this morning: a fresh reset was sitting on my account. Sam Altman’s post has the details. If there’s something you’ve been wanting to run full bore on the model you already have, this long weekend is a great time to do it. |
A weekly briefing for business leaders who want the few AI developments that matter.