Frontier AI models do not consistently seek out safety information before making a deployment decision, even when the stated chance of a problem rises sharply. That is the finding of SAFE, a controlled benchmark focused on evidence gathering before action rather than responses to warnings already placed in context.

Researchers tested GPT-5.5, o3, Claude Opus 4.8 and Claude Sonnet 4.6. The optional evidence varied in retrieval cost, probability, severity and presentation. Opus inspected evidence in most situations, while o3 skipped it most often and responded strongly to decision thresholds. GPT-5.5 and Sonnet fell between those patterns.

All models inspected more as potential harm became severe and inspected less as retrieving information became costly. Probability had a weaker effect: increasing the stated likelihood of a problem from 10% to 70% changed inspection rates by no more than 21 percentage points. That conflicts with the models’ written rationales, which frequently emphasized expected-value reasoning.

Further tests found that framing could change decisions near the inspection boundary even when models did not mention it in their explanations. The benchmark does not establish how agents will behave in open-ended deployments. It does show why safety evaluations should measure whether a system actively looks for missing evidence, not only whether it obeys a warning once shown.