The important safety failure in Anthropic’s new risk report was not a clever jailbreak. It was a whole model-development platform that sat outside the safeguard.
From May 2025 until April 2026, every exchange on Anthropic’s human-feedback contractor systems ran without its blocking biological-risk classifiers. Anthropic’s August 14 report puts the scale at roughly 50,000 people and 133 million exchanges. Most of those contractors could hold open-ended model conversations, rather than only rate a fixed answer.
These classifiers are meant to stop conversations that could materially help someone develop a biological weapon. But the failure was broader than leaving a filter switched off. A flag intended for internal use disabled both blocking and the logging of classifier alerts. Harmful-looking traffic therefore would not have been sent into Anthropic’s normal review process.
Anthropic’s retrospective review found no clearly concerning misuse. The company retained transcripts for almost all the affected traffic, except unsubmitted conversations on one platform representing about 1% of the total. It later ran Claude Sonnet 5 over every retained human turn and marked 1,197 transcripts as high risk. Of those, 757 came from Anthropic’s own teams. All but 62 of the rest came from deliberate red-team exercises.
Staff manually reviewed all 62 non-red-team transcripts and a random sample of 30 external red-team transcripts. They found a handful of potentially dual-use or academic red-team conversations, but nothing they judged to have given meaningful help to a threat actor.
That is real evidence against actual misuse. It is not proof that nothing happened: the review was retrospective, used a model prompt on isolated human turns, and could not recreate the real-time blocking and escalation that never occurred. Anthropic’s own conclusion is appropriately narrower. It says meaningful harm from this gap was very unlikely, while the discovery makes other unknown coverage gaps more likely.
The systems lesson is simple: a strong filter does not protect a route that never calls it. Safety claims often focus on how hard a classifier is to bypass. This incident asks the prior question: has every way of reaching the model actually been placed behind the classifier, with alerts that someone can see?
The contractor platform was not a customer product, and Anthropic says no customer data, model weights, internal systems or core networks were exposed. But it was still a route to capable models. Contractors were vetted by outside vendors, many of which, Anthropic says, previously lacked screening strong enough to stop even the lower-resource actors in its biological-risk model.
The same report describes a separate incident on these platforms: a small number of vendor contractors exploited a flaw to obtain an API key and use models outside their assigned work. The path existed for several weeks and included roughly two weeks of access to Mythos Preview without the biological classifiers. Anthropic says it contained that activity within 90 minutes of the external report and found no significant biological risk in the associated traffic.
Anthropic has now revised its own record. Its February risk report did not consider human-feedback platforms as a risk route. The company now rates the relevant biological-weapons risk as “low but not negligible,” and says its February assessment should have been low rather than very low. This is separate from the report’s better-publicized change to its model-misalignment rating.
What matters next is not another headline classifier score. It is evidence that the whole path works together: contractor identity checks, credentials, model routing, blocking, alert logging, human review and tightly governed exceptions. Anthropic fixed this known gap. The remaining question is how completely it has looked for the next one. Source graph: Semble source collection