Researchers introduced HarmProfile, a benchmark designed to describe the kinds of harmful outputs frontier language models produce when they fail safety tests. The dataset contains more than 80,000 validated artifacts from 23 frontier LLMs across 13 model families.

Most safety evaluations treat harmful generation as a pass-or-fail attack outcome. HarmProfile instead examines the content, severity, and variation of the harmful material itself, creating what the authors call a model-level risk profile.

That distinction matters because two models may fail at similar rates while producing very different kinds of harmful content. A content-centered benchmark can help researchers compare not just whether safeguards break, but how they break.

The paper is a research proposal and benchmark, not proof that any single model is safe or unsafe in the real world. Its contribution is a more detailed measurement approach for teams trying to understand safety failures beyond aggregate scores.