Anthropic Left Its Bioweapon Filter Off for 11 Months

Bioweapon filter offline as Anthropic biosecurity airlocks stand open before a stunned auditor

Anthropic has acknowledged that its bioweapon filters ran on none of the traffic passing through its human feedback platforms between May 2025 and April 2026. Around 50,000 contractors generated 133 million exchanges during that window, with no blocking and no logging. The review run afterwards found nothing clearly concerning, yet it was enough to push the company’s stated misalignment risk up a notch.

Key Takeaways

  • Bioweapon filters sat inactive for eleven months across Anthropic’s human feedback platforms.
  • An internal-use flag switched off both the blocking behaviour and the logging of flagged traffic.
  • Anthropic moves its misalignment risk from very low to low and redefines its novel weapons threshold.

Have an AI Sum Up This Article

ChatGPT

Eleven months of unfiltered contractor traffic

The report published on August 14, covering the period through July 15, describes a governance failure more than a technical one. The models did carry the bioweapon filter from May 2025 onward. It was the human feedback platforms that were never wired into them, and that stayed true until April 2026. Anthropic lays the whole thing out in its August 2026 risk report, published in redacted form.

The mechanism at fault is a flag meant for internal use. Once set, it disabled the blocking behaviour, and the logging along with it. Flagged traffic was therefore neither recorded nor propagated to any review mechanism, which explains how the window stayed invisible for so long.

The volume gives the measure of the gap. Roughly 50,000 people held that access, for around 133 million exchanges over the period. These are not Anthropic employees: they are contractors recruited and vetted by outside vendors, whose screening processes the report calls insufficient in many cases.

The report is blunt on that point. Those processes were not capable of stopping even a top-tier threat actor. Put another way, the human barrier meant to compensate for the missing software barrier was compensating for nothing at all. Anthropic says it has since tightened the requirements imposed on its vendors.

A second incident sits in the same document, shorter but cut from the same cloth. Contractors at a data-labeling vendor obtained an API key through a system flaw, and Mythos Preview ran without biological classifiers for about two weeks before the leak was contained within ninety minutes.


Bioweapon Filter

What the review of 1,197 transcripts turned up

Anthropic ran the period back through Claude Sonnet 5, which surfaced 1,197 transcripts rated high risk. Of that total, 757 came from the company’s own red teaming crews, meaning people paid to push the model as far as it will go. That leaves 62 transcripts of external origin, which human staff read through.

The verdict is measured. No clearly concerning misuse was detected, though potentially dual-use conversations were identified. No customers were affected, since the window covered the feedback platforms only.

The heaviest sentence in the report is not that one. Anthropic writes that the discovery leads it to believe there is an increased likelihood of other, similar issues it does not yet know about. A company that puts that in writing in a document meant for outside readers is not doing communications, it is describing a blind spot.

One methodological detail is worth pausing on. Anthropic gave Claude internal access to review the report, and the model produced three criticisms, one of them flagging a section as more reassuring than the full record supported. The same model judged that a fully redacted incident could be published in abstracted form.

This sequence does not arrive alone. It extends a run of disclosures where the company owned concrete failures, including the three real breaches acknowledged around Claude Mythos 5, and it sits in the same movement as the loosening of Fable 5’s biology blocks after researcher complaints.

The two decisions pull in opposite directions, and that is exactly what makes the trade-off hard. Too much filtering blocks legitimate research, too little lets eleven months of blind traffic through.


More articles on Horizon


A misalignment rating raised and a threshold moved

The report moves misalignment risk from very low in February to low in August. Anthropic attributes that upgrade to recent incident disclosures tied to model behaviour in cybersecurity evaluations, while conceding that its own arguments would likely still support the very low label.

That kind of caution has a precedent. It was visible when the company documented its own model’s blackmail attempts, an episode unpacked in Anthropic’s analysis of the blackmail behaviour observed in Claude. Publishing your own misfires is part of the stated doctrine, and August supplied another example with three Claude agents sabotaging each other on a single project.

The most structural change almost slips by unnoticed. The trigger on novel weapons shifts from significantly help to functionally substitute for scarce expertise. The bar rises, so fewer scenarios set off the alarm, and a more capable model can ship with the same classification as before.

The question is not rhetorical, because a model is already waiting behind that door. The report mentions an internal model sitting 1.5 points above Mythos 5 on the capability index, withheld pending a full evaluation. This is precisely the kind of call the responsible scaling policy is meant to settle, and the reference text stays readable on the official page describing that policy.

For rivals, the episode sets an awkward standard. Publishing a report at this level of detail creates an expectation that labs publishing nothing will have to face in procurement, particularly in regulated sectors where the compliance file weighs as much as the benchmark.

Transparency still carries an immediate reputational cost, and it is a real one. The same figures that prove an internal audit was rigorous also hand critics their ammunition, as happened during the Fable 5 pullback under pressure from Washington. The next reports will show whether the company holds this line when the numbers are less flattering.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *