Safety refusals have become a standard layer in consumer AI, but they are neither a complete technical defense nor a neutral rulebook. MIT Technology Review examines how model providers train systems to reject requests they classify as dangerous while preserving the underlying capabilities.

Early language models would answer almost any prompt. Today, companies reward models for refusing harmful tasks, penalize unnecessary refusals and place additional classifiers in front of them. The result reduces routine access to instructions for violence, cyberattacks or self-harm, but determined users still find ways around probabilistic controls.

The boundary also depends on context. A request about a pathogen or software vulnerability may support legitimate research, defense or abuse. Providers decide how to distinguish those cases, generally without disclosing the full policy or training process. That can leave users unable to tell whether a refusal reflects law, safety evidence, a commercial preference or a government demand.

The article’s central warning is not that models should answer every request. It is that refusal behavior should not substitute for limiting dangerous capability, controlling tool access and auditing consequential use. As models gain access to laboratories, networks and autonomous systems, reliable containment requires more than teaching the same capable model to say no.