Extraordinary and dangerous. But is it real?

The labs’ word is no longer enough. Here’s what you should verify yourself.

The number of announcements of new models and the damaging potential of those models has really been bothering me.

After we had another breach announcement this week, I took a little bit more time to dig into these breaches and see if I can figure out what's going on behind the scenes.

These models are incredibly capable, and when backed into a corner, they'll do almost anything to succeed on their mission, even if that includes breaking things to make it happen.

Also, a couple quick announcements on the release of Fable 5.1 and GPT-6 Astra. These are highly capable models, but they're only as powerful as the access and context that you give them, so getting clear on the guardrails and your supporting folder structure become more and more important.


01 — The sandboxes aren't working

What happened

If I've learned anything about sand over my 41 years on this planet, it's that it never stays where it's supposed to. It always finds a way out of the sandbox and into everything.

Seems like the pattern's repeating itself at the frontier AI companies.

Major frontier models at both OpenAI and Anthropic escaped test environments this summer and reached other people’s systems.

  • July: about 700 OpenAI agents cheated on a cyber benchmark, got into Hugging Face’s systems, and tried to cover their tracks. Hugging Face disclosed July 16. OpenAI took responsibility July 21 and published a report August 26. Fortune noted what the report leaves out: the prompt the agents were given.
  • After that news, Anthropic reviewed 141,006 of its own evaluation runs. Three Claude models had reached the internet from sandboxes meant to be sealed, then got into production systems at three organizations through weak passwords and endpoints with no login. One published a malicious Python package that was downloaded and run on 15 real machines.
  • Tuesday: Anthropic disclosed that by April, more than 10% of its production training environments had been flagged for reward hacking. Its response: about 150 product engineers moved to security work, outbound traffic blocked by default on compute clusters, a real-time classifier watching for escape attempts, outside cyber evaluations paused and rebuilt, and METR brought in to review.
  • Same week, both labs shipped anyway. Anthropic released Claude Fable 5.1 on Tuesday; the unrestricted version, Mythos 5.1, goes only to vetted cyber and life-sciences organizations.
  • OpenAI began a limited release of GPT-6 Astra on Thursday to its Daybreak program for cybersecurity defenders, calling it the first model to trigger its advanced internal safety protections for cyber capabilities. Per NBC, OpenAI also disclosed that an unreleased model in Astra’s family had, on its own, taken administrator control of part of OpenAI’s own infrastructure.

What changes

Neither lab caught it first. The victim disclosed before OpenAI did, Anthropic went looking after the OpenAI news, and the April incident they uncovered took five months to reach the public. You can no longer rely on the labs’ own monitoring.

But why do the models escape? The OpenAI agents decided the tasks were impossible and cheated to produce an answer anyway. These systems are trained to please, to accomplish their mission, and when the test can’t be passed (or they think they can't be) they break the test rather than come back empty-handed.

Anthropic’s own researchers showed the mechanism on purpose: train a model on 80 environments with known exploits and it generalizes, unprompted, to breaking out of sandboxes, stealing credentials, and attacking infrastructure to grab the answer key. The models in your stack came out of the same kind of training.

Notice the shape of the week, too. The containment failures and the 10% figure came out Tuesday. By Thursday, both labs had shipped their strongest models, with the most capable versions held back for approved customers.

“Too dangerous to release” is a convenient conclusion when you’re the one who already has the model. Every restriction on who gets the most capable version is also a moat.

This leads into the ongoing debate between closed and open-source models and the need for regulation. But that's a topic that we'll surely cover in another issue.

What to do next

  • Scope each agent’s access to what its step needs, and nothing more. Anthropic’s fix was to block outbound traffic by default on its compute clusters. Yours is narrower: no internet, no send, no delete unless that step requires it.
  • Where one job needs several kinds of access, split it into several agents. My email triage is built on that split, and this is a real example from my daily use:
    1. The agent that reads the inbox can label but cannot draft or send.
    2. The agent that reviews labels and drafts messages cannot send them.
    3. Only I am able to pass final judgment on a message and send it.
    The step that reads untrusted email is never the step that can act on it.
  • Give each agent its own credentials, scoped to that step, and rotate them. If an agent has access to something, it eventually will find it. Security by obscurity is dead in the AI era.
  • Log agent activity somewhere the agent can’t write. Agents can and will manipulate outputs and tests if they know they're being tested against them. You need strict guardrails in place and, ideally, immutable storage to hold the results.
  • If you want to run sensitive or proprietary work through Claude, look at Enterprise Frontier Safeguards, announced Tuesday: your 30-day misuse-monitoring logs stay in your own S3, Azure Blob, or GCS storage instead of Anthropic’s servers, with zero data retention. Phased rollout this fall.

Anthropic’s post · OpenAI’s report · Fortune on what OpenAI’s report leaves out · METR and Redwood Research’s independent investigation · Redwood on why the agents weren’t just following instructions · The Register on the Anthropic incidents · Fable 5.1 and Mythos 5.1 · NBC on Astra · Enterprise Frontier Safeguards · Open Secure AI Alliance for AI Safety and Security


02 — The EU asks for proof

What happened

  • August 29: the EU AI Office sent formal requests for information to general-purpose model providers. First use of enforcement powers that took effect August 2.
  • Euractiv reports more than 30 providers, including OpenAI, Anthropic, and Google.
  • Incomplete or misleading reply: up to €15 million or 3% of global turnover, and the Commission can restrict a model’s availability in the EU.
  • Per reporting, four topics: security against attack, independent evaluation, post-release monitoring, training data.

What changes

The AI Act is now a questionnaire with a deadline and a fine for misleading answers.

The four topics cover many of the things the labs just publicly admitted they were getting wrong, and the answers going to Brussels will be more careful than anything published voluntarily.

What to do next

Ask these four questions about your current model vendor:

  1. How is the model protected against attack?
  2. Who evaluated it independently, and can I see the results?
  3. How do you monitor it after release, and what has that turned up?
  4. What was the model trained on?

Or skip the rebuild:

Send me what you sent the AI Office.

What you find out will certainly beat what you know now. This is vital for anyone serving EU customers, and good practice for the rest of us.

The Commission’s enforcement page


03 — OpenAI says, "No Model for You, Musk."

What happened

OpenAI is cutting Cursor off from its models on November 12, 75 days after telling them.

  • SpaceX bought Cursor for a reported $60 billion, closing mid-August. Two weeks later OpenAI said it will wind down access.
  • OpenAI’s stated reason: its experience after Elon Musk’s acquisition of Twitter, and Musk’s sworn testimony that xAI had violated OpenAI’s terms.
  • Cursor triggered nothing. Its new parent company’s founder has a history with OpenAI, and that was enough.
  • Cursor CEO Michael Truell says OpenAI models serve about 5% of its customers, and Anthropic said in the same news cycle it would keep adding compute for Claude in Cursor. Cursor has leverage you don’t: a public CEO, a competitor offering help, and 75 days of notice. Your notice will be an email, if you get one.

What changes

You don’t own your access. You rent it. Amazon showed the same thing in 2009 when it deleted paid copies of 1984 from customers’ Kindles over a rights dispute.

We are beholden to the systems and what they provide for us.

What to do next

  • Find the one model you couldn’t move off of in 75 days. Fix that dependency first.
  • Run a real workflow on a second provider’s model and check the output. Only a tested fallback helps you in November.
  • Keep your prompts, skills, and evals where you can take them with you. In your repo, in plain files. Anything that lives only inside a vendor’s product leaves with the vendor.

OpenAI’s post


Also worth knowing: Better tests are coming

Google DeepMind ran the first double-blind evaluation of a proprietary frontier model. Evaluators can’t see the weights; Google can’t see the test questions. One small model, one pilot.

If it becomes normal, “our model scored X” starts to mean something again. DeepMind’s post


One question to close on. How much of your AI setup could you copy to a new vendor in an afternoon?

Hit reply and tell me the number. I read every one.

–KM

P.S. While Astra rolls out, OpenAI is banking one full usage reset per day for every paid ChatGPT account that doesn’t have Astra access yet. I confirmed it this morning: a fresh reset was sitting on my account.

Sam Altman’s post has the details. If there’s something you’ve been wanting to run full bore on the model you already have, this long weekend is a great time to do it.

Worth Your Attention

A weekly briefing for business leaders who want the few AI developments that matter.