Anthropic's Claude Fable 5 topped six new industry-specific benchmarks from Artificial Analysis, including tests for finance, law, and medicine. The results came with a notable cost premium.
Vertical benchmarks are becoming more important as buyers ask whether frontier models can handle specialized professional work, not just general chat or coding tasks. But per-task pricing can change the practical value of benchmark leadership.
The report underlines a key enterprise tradeoff: the best model for accuracy may not be the best model for routine production volume.