Anthropic disclosed a serious safety-control gap: internal classifiers meant to block biological and chemical weapons risk were inactive for external contractor feedback traffic from May 2025 through April 2026.

The filters are designed to prevent models from helping users extract dangerous knowledge about chemical or biological weapons. During the outage, about 50,000 external contractors ran roughly 133 million interactions with Anthropic models without those classifiers applied.

Anthropic says its investigation found no evidence of actual misuse. The company also said it has tightened contractor requirements after finding that vendor screening processes were often insufficient.

The disclosure is notable because Anthropic’s leadership has repeatedly emphasized biosecurity as a major AI risk. It also shows how safety depends on deployment plumbing, not only model behavior. A classifier can be well designed and still fail to protect traffic if the operational path around it is misconfigured or excluded.