OpenAI has introduced LifeSciBench, a benchmark for evaluating AI systems on real-world life-science research tasks and decisions. The benchmark is expert-authored and expert-reviewed.

The launch matters because scientific AI systems need evaluation beyond general reasoning tests. Life-science work involves domain-specific constraints, experimental judgment, and decisions where mistakes can be costly.

Benchmarks like this can help researchers compare models on tasks closer to actual scientific workflows, rather than relying only on broad academic exams.