Skip to main content
Back to Blog
AI
Security
Risk
Governance
Enterprise AI
OpenAI

One in Twelve Still Gets Through.

Jason Oglesby

By Jason Oglesby · September 5, 2026

On September 1, OpenAI wrote that its new model can "find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step."

That is their sentence, not mine. Astra is the first model they have designated at the Critical cybersecurity level under their own Preparedness Framework.

Then they published a number. Jailbreak refusal went from 59 percent on the previous model to 91.5 percent on this one.

Disclosures. I pay OpenAI, Anthropic, xAI and Perplexity, so the company I am about to examine takes my money and so do three of its competitors. And Ergon sells technology leadership work that includes security posture, so a post about your security assumptions is a post that sells what I do.

They Said It First, and That Counts

I complained on Thursday that OpenAI's capability claims arrived without a single number attached. This week they published the framework, the threshold, the designation, and the evaluation results, before shipping rather than after being caught.

That is not nothing. Most vendors disclose this class of finding when a researcher forces them to. OpenAI wrote it down first, in their own words, on their own site, and used language that makes their product sound dangerous. Give them the credit.

Which is what makes the number worth taking seriously instead of dismissing.

Invert It

91.5 percent up from 59 percent is a real improvement. It is also 8.5 percent.

Roughly one attempt in twelve got past the model's refusal training in their own evaluation.

Be careful about what that does and does not mean. It is a refusal rate against jailbreak attempts in a controlled test, not a breach rate in the world. And refusal is one layer of several. OpenAI also describes system-level classifiers, activation detectors, chain-of-thought monitoring for misaligned behavior, tighter boundaries on high-risk accounts, and a round-the-clock red-team program. Access to the advanced cyber workflows starts with a small alpha group and widens through a defender program.

They built those other layers because they know a percentage is not a wall. That is the correct engineering response.

The mistake is not theirs. It is the one a buyer makes reading the headline number as the safety story when the vendor themselves treats it as one control out of six.

Percentages Are the Wrong Unit Here

A 91.5 percent refusal rate is excellent against a curious user who tries once and moves on.

It means something different against someone who is patient, funded, and willing to try continuously. For that adversary a refusal rate is not a barrier, it is a cost of doing business, and the arithmetic runs in their favor over enough attempts.

This is why the other layers matter more than the headline figure, and why the number I would actually want is not the refusal rate. It is how quickly the monitoring catches a pattern of attempts, and what happens to that account when it does.

That number is not published. I am not saying it is bad. I am saying we do not have it, and it is the one that governs the outcome.

What Changes on Your Side

Most security programs are priced against an assumption that has been true for thirty years: finding a novel vulnerability in a hardened system is slow, expensive, and requires a scarce human.

Your patch cadence assumes it. Your risk register assumes it. Your decision to defer that upgrade another quarter assumes it.

A model designated Critical for exactly that work is a change to the assumption, not to the tooling. Even under restricted access. Even used only by defenders, which is genuinely how it starts. Capability that exists gets cheaper and more available on a schedule nobody controls.

The right response is not alarm. It is looking at which of your decisions were priced on the old assumption.

What I'd Do This Week

Find the deferrals. Every "we will patch that next quarter" was a bet on how hard the thing is to exploit. Reprice the bet, not the patch list.

Ask what your monitoring would catch. Not whether you would block an attempt. Whether you would notice a thousand of them.

Stop treating a refusal rate as a control. It is a property of a model. A control is something you own and can test.

Read your vendors' safety documentation, not their launch posts. OpenAI published theirs. Most of what matters this week was in it, and almost none of it was in the coverage.

The Part That Matters

The company that built the thing told you what it can do, in plain language, before shipping it. That is the most useful thing that happened this week.

The number they published is good and getting better. One in twelve is also one in twelve.

Both of those are true, and only one of them is in your threat model.