The Defense That Concedes the Jailbreak: Decoy Hardening Poisons What Removal Unlocks

TV
Thiago Victorino
8 min read
The Defense That Concedes the Jailbreak: Decoy Hardening Poisons What Removal Unlocks

In August 2026, Mark Russinovich of Microsoft Azure published a defense built on a concession: abliteration strips a model’s refusal mechanisms in minutes on consumer hardware. Fool’s Gold does not try to stop the strip. It trains the model so that what the attacker unlocks is poisoned. The same month, OpenAI paused reinforcement learning training on its deployment-bound models for two weeks and put its largest planned frontier RL run on hold, after a model called Astra showed preliminary evidence of Critical cybersecurity capability.

Two moves, same inversion. One vendor stopped defending the lock and started degrading the loot. One lab stopped gating deployment and started throttling training itself. Prevention is no longer the operating assumption of frontier security governance. Degrading the attacker’s payoff is.

A Defense Designed for the Post-Removal State

We covered the attack side of this in the Qwen censorship-circuit analysis: alignment sits on top of a model like a decal, and removal is cheap. Fool’s Gold is the first published defense that takes that finding as its starting condition rather than its threat model. The design question changes from “how do we keep refusals in place” to “what does the model do for the attacker once refusals are gone.”

The answer is confident nonsense, selectively. A defended model, after its safety training is stripped, produces decoy answers on harmful requests at rates of 0.51 to 0.90 across the six models tested. The decoys are not visible refusals or degraded gibberish. They read as competent answers and are wrong where wrongness matters: on CBRNE-adjacent benchmarks, the defended and then jailbroken model is fatally wrong on 0.82 to 0.86 of answers, against at most 0.10 for an undefended model under the same attack.

The two properties that make this a governance artifact rather than a lab curiosity are the clean-state numbers. First, refusal behavior before any attack stays at 0.997 to 1.000, so the defense does not leak into normal operation. Second, benign capability on MMLU, GSM8K, WMDP-bio, and IFEval stays within noise. The legitimate user pays nothing measurable. The attacker inherits a model that answers everything and can no longer be trusted on precisely the questions the jailbreak was for.

Russinovich’s framing sentence carries the whole design: “What cannot be prevented can be deceived.”

An Open Recipe, a Gated Corpus

The release is not a paper plus a demo. The corpus-construction pipeline and the simulated-attack training harness are open-source and config-driven, described as one JSON file per model. Any lab or vendor shipping open-weight models can run the recipe against its own checkpoints. The real decoy corpus is gated to verified researchers, which is the sensible split: the method spreads, the ammunition does not.

That distribution choice matters for the vendor conversation. Until now, an open-weight vendor asked about safety removal had one honest answer: it happens, and we cannot stop it. The gated-model line of defense responds by not shipping weights at all. Decoy hardening gives the open-weight side its first affirmative answer: removal still happens, and what it unlocks is booby-trapped. Whether that answer holds up under adaptive attacks designed to detect decoys is an open question. But a defense with published rates, published costs, and a runnable pipeline is now the benchmark an open-weight vendor can be measured against.

The Market Already Prices the Same Assumption

The offense side reached the same conclusion from the other direction. Claudia d’Antoine, CEO of Margin Research, describing AI-assisted vulnerability research in August 2026: “Before an offensive researcher is even able to find a vulnerability or exploit it, it is already halfway to n.” Her firm calls the resulting artifact the half-day: an exploit whose exclusivity window collapses when AI-driven scrutiny is already halfway to the same flaw.

A half-day market is a market that has stopped assuming secrets keep. That is the economic mirror of decoy hardening. If you cannot keep the vulnerability exclusive, the value shifts to speed and to what the target does after compromise. If you cannot keep the refusal in place, the value shifts to what the model does after removal. We traced the offensive side of this compression in how AI offense rewrites open source. The defensive vocabulary is now catching up to it.

The Lab Throttles Its Own Input

OpenAI’s pacing announcement is the third convergent data point, and the most institutionally expensive one. Per the post, which we read directly: on August 7 the company determined that Astra, a frontier model in training, showed preliminary evidence of Critical cybersecurity capability under its Preparedness Framework. Combined with a security incident involving Hugging Face, whose details are not public and whose technical report is promised in the coming weeks, the response was a two-week pause in RL training on deployment-bound models and a hold on the largest planned frontier RL run.

Pausing training is a different category of control than gating access. The trusted-access program governs who reaches a capability after it exists. A training pause governs whether the capability comes to exist on schedule. The control moved upstream of the artifact.

The announced operating numbers show what running that posture costs. The post commits to an alert SLA of 30 minutes and puts monitoring overhead at roughly 20% of the inference compute being monitored. One fifth of monitored inference spent watching the other four fifths is a real line item, disclosed voluntarily, by the party paying it. And the post’s forward-looking sentence closes the loop with the other two sources: “We expect models to soon drive most security work, including defending against other models.”

What the Inversion Changes for Buyers

If you procure or deploy open-weight models, the diligence question has a new shape. Asking a vendor whether their safety training can be removed is now a question with a known answer, and a vendor who claims otherwise is telling you something about their threat model, none of it good. The productive question is what the model yields to the attacker in the post-removal state, and whether the vendor can put numbers on it the way Fool’s Gold does: decoy rate, clean-state refusal, benign-capability delta, fatal-error rate on the domains that matter.

If you run internal fine-tunes of open-weight models, the same recipe cuts the other way. A hardened base model that emits confident decoys after safety removal will emit those decoys to your red team too, and possibly to a fine-tune that accidentally disturbs the safety training. Decoy behavior needs to be in your evaluation matrix before a hardened model enters your stack, or you will be debugging plausible wrong answers with no signature to grep for.

Do this now: add a post-removal column to your model evaluation sheet. For every open-weight model in use or under consideration, record what is known about its behavior after safety-training removal: nothing, vendor claims, or measured rates. Today almost every row will say nothing. That column is where this class of defense will be scored from now on, and the first vendor conversations that reference decoy rates are the ones your security team should be in the room for.


This analysis synthesizes Fool’s Gold: Defensive Deception Against Safety-Removal Attacks (Mark Russinovich, Microsoft Azure, August 2026), Introducing the Half-Day: 0-Day in the Age of AI (Margin Research, August 2026), and Pacing Model Development in an Era of Cyber-Critical Capabilities (OpenAI, August 2026).

Victorino Group helps engineering organizations evaluate open-weight model risk, including post-removal behavior, before it enters the production stack. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation