This article is a deep-dive from JudyAI Lab — an AI engineering playbook series with 100+ published guides, 5,000+ weekly readers across 60+ countries, focused on the practical side of running AI agents, trading systems, and content pipelines in production.
ServiceNow AI's research team has released EVA-Bench Data 2.0, an enterprise-grade benchmark designed specifically for voice agents. This release dramatically expands the scope, moving from a single domain to three major enterprise scenarios: airline customer service management (CSM), enterprise IT service management (ITSM), and healthcare human resources service delivery (HRSD). Together, the three domains cover 213 evaluation scenarios and 121 tools, roughly four times the coverage of the original version. Broken down by domain: airline has 50 scenarios, ITSM has 80, and HRSD has 83.
This benchmark places particular emphasis on real-world voice scenarios — every data point was filtered starting from actual phone-based customer service workflows, with tool schemas modeled after production API specs. The healthcare HRSD domain goes even deeper, tying into real US healthcare policy details like NPI provider identifiers, FMLA family leave regulations, and insurance coverage rules, ensuring the evaluation scenarios match what practitioners actually deal with day to day. All 213 scenarios were cross-validated for solvability by three frontier models — OpenAI's GPT-5.4, Google's Gemini 3.1 Pro, and Anthropic's Claude Opus 4.6 — to keep the benchmark challenging while ensuring results stay fair and trustworthy.







