ServiceNow Research published EVA-Bench Data 2.0, significantly expanding the benchmark for evaluating enterprise AI agents. The new version covers three business domains—CRM, HR, and IT—with 121 distinct tools and 213 evaluation scenarios drawn from real enterprise workflows.
The benchmark is designed to test how well AI agents handle multi-step tasks that require tool use, data retrieval, and contextual reasoning across business systems. Each scenario mirrors actual enterprise operations such as updating customer records, processing HR requests, and resolving IT tickets.
EVA-Bench has become a key reference point for the enterprise agent community. The 2.0 release adds more complex, interleaved tool-use patterns and richer ground-truth annotations, making it harder for agents to succeed through simple pattern matching. The dataset is available on Hugging Face under an open license.