0% found this document useful (0 votes)
7 views7 pages

Clear

The document critiques current benchmarks for evaluating enterprise agentic AI systems, highlighting their focus on accuracy while neglecting cost, reliability, and operational stability. It introduces the CLEAR framework, which incorporates multiple dimensions such as cost, latency, efficacy, assurance, and reliability to provide a more comprehensive evaluation for enterprise applications. The findings demonstrate that optimizing solely for accuracy can lead to significantly higher costs and unreliable performance, emphasizing the need for a multidimensional approach to agent evaluation.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views7 pages

Clear

The document critiques current benchmarks for evaluating enterprise agentic AI systems, highlighting their focus on accuracy while neglecting cost, reliability, and operational stability. It introduces the CLEAR framework, which incorporates multiple dimensions such as cost, latency, efficacy, assurance, and reliability to provide a more comprehensive evaluation for enterprise applications. The findings demonstrate that optimizing solely for accuracy can lead to significantly higher costs and unreliable performance, emphasizing the need for a multidimensional approach to agent evaluation.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise

Agentic AI Systems

Sushant Mehta
sushant0523@[Link]

Abstract First, cost is entirely ignored. Despite agents making


arXiv:2511.14136v1 [[Link]] 18 Nov 2025

hundreds of API calls per task, with complex architectures


Current agentic AI benchmarks predominantly evaluate task
completion accuracy, while overlooking critical enterprise re- like Reflexion making up to 2,000 API calls for iterative re-
quirements such as cost-efficiency, reliability, and operational finement, no major benchmark reports cost metrics. Our
stability. Through systematic analysis of 12 main bench- analysis shows that leading agents exhibit 50x cost varia-
marks and empirical evaluation of state-of-the-art agents, we tions (from $0.10 to $5.00 per task) for similar accuracy
identify three fundamental limitations: (1) absence of cost- levels, with complex agent architectures achieving marginal
controlled evaluation leading to 50x cost variations for sim- accuracy gains at exponential cost increases (Kapoor et al.
ilar precision, (2) inadequate reliability assessment where 2024). This creates a distorted research landscape in which
agent performance drops from 60% (single run) to 25% (8- expensive and fragile solutions appear superior to efficient
run consistency), and (3) missing multidimensional metrics alternatives. A 2-point improvement in accuracy might cost
for security, latency, and policy compliance. We propose
$50,000 additional spend per 10,000 tasks, an unacceptable
CLEAR (Cost, Latency, Efficacy, Assurance, Reliability), a
holistic evaluation framework specifically designed for enter- trade-off for most enterprises.
prise deployment. Evaluation of six leading agents on 300 en- Second, reliability remains unmeasured. Production
terprise tasks demonstrates that optimizing for accuracy alone systems require consistent performance across thousands of
yields agents 4.4-10.8x more expensive than cost-aware al- similar requests with failure rates below 1-5%, yet bench-
ternatives with comparable performance. Expert evaluation marks report single-run success rates that mask brittleness.
(N=15) confirms that CLEAR better predicts production suc- Recent work in the τ -bench (Yao et al. 2024) introduced
cess (correlation ρ = 0.83) compared to accuracy-only eval- pass@k metrics, revealing that GPT-4-based agents drop
uation (ρ = 0.41). from success 60% (pass@1) to just 25% (pass@8), which
is insufficient for enterprise deployment where reliability is
1 Introduction paramount. A customer service agent that works 70% of the
Rapid advancement of autonomous agents based on large time in testing but only 30% consistently in production cre-
language models (LLM) has generated significant interest ates a poorer user experience than a 60% agent with reliable
in their enterprise applications, from software engineering performance.
(Jimenez et al. 2024) to customer service automation (Yao Third, enterprise-critical dimensions are absent. Real-
et al. 2024). However, there is a critical gap between world deployment demands security against prompt injec-
benchmark performance and production deployment suc- tion, compliance with organizational policies, latency within
cess. While 85% of the companies experiment with gener- SLA constraints, and graceful error handling, none system-
ative AI, only a small fraction deploy agents in production, atically evaluated by current benchmarks. Current bench-
with most projects abandoned after proof-of-concept stages marks evaluate none of these, creating a performance gap
(Chiu 2025). This failure stems from a fundamental mis- 37% between lab tests and production deployment (Liu et
alignment: existing benchmarks optimize for task comple- al. 2024). An agent that passes all functional tests, but leaks
tion accuracy, while enterprises require holistic evaluation customer data or violates regulatory requirements, repre-
across cost, reliability, security, and operational constraints. sents a catastrophic failure, not a minor shortcoming.
Recent benchmarks have made substantial progress in This work makes four primary contributions:
evaluating agent capabilities. SWE-bench (Jimenez et al. (1) Systematic Gap Analysis: We analyze 12 major
2024) assesses software engineering through real GitHub agentic benchmarks (SWE-bench, WebArena, AgentBench,
issues, WebArena (Zhou et al. 2023) tests web navigation GAIA, ToolLLM, τ -bench, WorkArena, Mind2Web, OS-
across realistic environments, and AgentBench (Liu et al. World, InterCode, BFCL, and others) and identify specific
2023a) provides multi-environment evaluation. However, limitations through the lens of enterprise requirements, doc-
systematic analysis reveals three critical limitations that im- umenting validity issues affecting 7/10 benchmarks and cost
pede enterprise adoption and create a distorted view of agent misestimation rates up to 100% (Kang et al. 2024).
capabilities. (2) CLEAR Framework: We propose a five-dimensional
holistic evaluation framework (cost, latency, efficiency, on ServiceNow with 33 atomic tasks, while τ -bench (Yao et
assurance, reliability) with novel metrics including cost- al. 2024) introduces multi-turn user interactions with policy
normalized accuracy (CNA), pass@k reliability, policy ad- compliance in retail and airline domains, revealing reliabil-
herence score (PAS), and SLA compliance rate tailored for ity issues through pass@k metrics.
enterprise contexts. The framework supports both individual
dimension analysis and composite scoring with customiz- 2.2 Benchmark Limitations
able weights. Recent critical analyses expose systematic problems.
(3) Enterprise Task Suite: We introduce a bench- Kapoor et al. (Kapoor et al. 2024) demonstrate that agent
mark of 300 tasks across six enterprise domains (customer evaluations fail to control cost, creating misleading con-
support, data analysis, process automation, software de- clusions about architecture improvements. Their analysis
velopment, compliance, multi-stakeholder workflows) with shows that simple baseline strategies outperform complex
ground-truth cost, latency, and policy compliance annota- agents at 50x lower cost. They identify inadequate hold-
tions. Each task includes 5-15 steps with realistic complex- out sets, with 7/8 benchmarks lacking appropriate validation
ity that reflect actual enterprise workflows. sets for claimed generality levels.
(4) Empirical Validation: Evaluation of six state-of-the- Kang et al. (Kang et al. 2024) found severe validity issues
art agents (ReAct-GPT4, ReAct-GPT-o3, Reflexion, Plan- in 8/10 popular benchmarks, including task validity failures
Execute, ToolFormer, Domain-Tuned) reveals that accuracy- (do-nothing agents passing 38% of τ -bench airline tasks),
optimal configurations cost 4.4-10.8x more than Pareto- outcome validity failures (LLM-as-judge making arithmetic
efficient alternatives. Expert validation (N = 15 Enterprise- errors, augmented tests changing 41% of SWE-bench rank-
AI leads) shows that CLEAR predictions are strongly corre- ings), and undisclosed issues. Yehudai et al. (Yehudai et
lated with production success (ρ = 0.83, p<0.001) versus al. 2025) surveyed 120 agent evaluation frameworks, iden-
accuracy-alone (ρ = 0.41, p=0.03), establishing statistical tifying missing enterprise requirements including multistep
significance and practical relevance. granular evaluation, cost-efficiency measurement, focus on
Our findings establish that multidimensional evaluation safety and compliance, and live adaptive benchmarks.
is not optional but essential for the deployment of enter-
prise agents, providing actionable frameworks for both re- 2.3 Enterprise AI Evaluation
searchers and practitioners. We aim to release our Enterprise
Task Suite, evaluation code, and complete experimental re- Enterprise evaluation differs fundamentally from academic
sults to facilitate adoption and enable reproducible research benchmarking. Aisera’s CLASSic framework (Aisera Re-
to advance practical agent systems. search Team 2024) proposes five dimensions (Cost, Latency,
Accuracy, Stability, Security) with empirical evidence that
2 Related Work domain-specific agents achieve 82.7% accuracy versus 59-
63% for general LLMs at 4.4-10.8x lower cost. AWS re-
2.1 Agentic AI Benchmarks search (Liu et al. 2024) on multi-agent systems demonstrates
The landscape of agent evaluation has evolved rapidly. Soft- 90% goal success rates with proper coordination versus 53-
ware engineering benchmarks include the SWE-bench 60% for single agents, while documenting a 37% perfor-
(Jimenez et al. 2024), which evaluates agents on 2,294 real mance gap from lab to production. Industry reports high-
GitHub issues with execution-based testing, and InterCode light that only 10% of enterprises successfully implement
(Yang et al. 2023) enabling interactive coding with execu- generative AI in production (Chiu 2025), with inadequate
tion feedback across Bash, SQL, Python, and CTF chal- evaluation frameworks cited as the main failure factor.
lenges. Web navigation benchmarks assess agents in
browser environments: WebArena (Zhou et al. 2023) pro- 3 The CLEAR Framework
vides 812 tasks on self-hosted websites with functional cor- We propose CLEAR (Cost, Latency, Efficacy, Assurance,
rectness evaluation, while Mind2Web (Deng et al. 2023) Reliability), a comprehensive framework that addresses the
covers 2,350 tasks on 137 real websites. identified gaps. CLEAR recognizes that enterprise deploy-
Multi-domain benchmarks test generalization. Agent- ment requires multi-objective optimization across five criti-
Bench (Liu et al. 2023a) evaluates 29 LLMs in eight en- cal dimensions.
vironments (OS, database, knowledge graphs, gaming, em-
bodied AI), revealing significant gaps between commercial 3.1 Framework Dimensions
and open source models. GAIA (Mialon et al. 2023) pro-
vides 466 real-world questions that require reasoning, mul- Cost (C) measures economic efficiency, including API to-
timodality, and tool use, exposing a 77% human-AI perfor- ken consumption, inference costs, and infrastructure over-
mance gap. Tool-use benchmarks include ToolLLM (Qin head. We introduce cost-normalized accuracy (CNA):
et al. 2024) that includes 16,464 real-world APIs from Rapi- Accuracy
dAPI and Berkeley Function Calling Leaderboard (Patil et CNA = × 100 (1)
al. 2024) that assess function calling across multiple lan- Cost
guages and scenarios. where Cost is measured in USD per task. This metric en-
Enterprise-focused benchmarks have emerged recently. ables a fair comparison between expensive high-accuracy
WorkArena (Drouin et al. 2024) evaluates knowledge work agents and cost-effective alternatives. Furthermore, we track
the cost per success (CPS) to account for varying success
rates: Table 1: Performance across CLEAR dimensions. Pareto-
Total Cost optimal agents marked with *.
CPS = (2)
Number of Successful Tasks Agent Eff. Cost CNA Lat. PAS R@8
CPS reveals that failed attempts still incur costs, making re- (%) ($) (s) (%)
liability economically critical. ReAct-GPT4 72.3 2.87 25.2 8.4 0.89 58.3
Latency (L) evaluates response time throughout the plan- ReAct-GPT-o3* 68.7 0.31 221.6 4.2 0.85 52.1
ning, execution, and reflection phases. We measure end- Reflexion 74.1 5.12 14.5 12.7 0.91 61.2
to-end task completion time and introduce SLA compliance Plan-Execute* 71.9 1.24 58.0 6.8 0.88 64.5
rate: ToolFormer 69.5 1.89 36.8 5.9 0.82 55.7
Domain-Tuned* 70.3 0.27 260.4 3.8 0.93 72.8
Tasks Completed Within SLA
SCR = × 100% (3)
Total Tasks
with domain-specific SLA thresholds (3 seconds for cus- We developed 300 tasks in six domains with ground-truth
tomer support, 30 seconds for code generation). Latency di- CLEAR annotations: Customer Support (60 tasks): Mul-
rectly impacts user experience and system throughput, mak- titurn policy-compliant issue resolution with escalation han-
ing it critical for production deployment. dling; Data Analysis (50): SQL query construction, report
Efficacy (E) captures task completion quality through tra- generation, visualization; Process Automation (50): Mul-
ditional accuracy metrics augmented with domain-specific tistep workflows with approval chains; Software Develop-
measurements. For software tasks, we evaluate functional ment (60): Fixing bugs, review of code, generation of test
correctness through test passage rates. For data analysis, we from production repositories; Compliance (40): GDPR pro-
assess the accuracy and quality of the result. For customer cessing, regulatory verification; Multi-Stakeholder (40):
support, we measure the accuracy of the intention classifica- Cross-departmental coordination with conflicting priorities.
tion and the appropriateness of the response. Unlike single- Tasks feature 5-15 steps with realistic complexity. See Ap-
metric benchmarks, we recognize that efficacy requirements pendix A for detailed descriptions.
vary by domain and use case.
Assurance (A) evaluates safety, security, and policy com- 4 Experimental Evaluation
pliance. We introduce policy adherence score (PAS):
Setup. We assessed six architectures: ReAct-GPT4, ReAct-
Policy Violations GPT-o3, Reflexion (Shinn et al. 2023), Plan-Execute (hi-
PAS = 1 − (4)
Total Policy-Critical Actions erarchical planner), ToolFormer (Schick et al. 2023), and
Security assessment includes prompt injection resistance Domain-Tuned (fine-tuned Llama). Each executed all 300
in 500 adversarial test cases from established attack tax- tasks. For reliability, we sampled 60 representative tasks and
onomies, data leak prevention, hallucination rates for do- executed each 10 times. We recruited 15 AI enterprise de-
main knowledge, and graceful failure handling. Policy vio- ployment leads (mean experience 5.9 years) who evaluated
lations represent hard failures in enterprise contexts: a single the results on 40 randomly assigned tasks, rating readiness
unauthorized data disclosure can invalidate otherwise per- to deploy on a 5-point scale (inter-rater reliability α = 0.78).
fect performance. Results. Table 1 presents results across CLEAR dimen-
Reliability (R) assesses consistency through the pass@k sions. Although Reflexion achieved the highest efficacy (74.
metric introduced by τ -bench (Yao et al. 2024), where 1%), it cost 5.12 times more than the efficacy of ReAct-
pass@k measures the probability of achieving k consecutive GPT-o3 (68. 7%, representing only 5.4 percentage points
successes: of improvement with the increase in cost 1, 551%. Domain-
Tuned achieved the best cost-normalized accuracy (260.4)
Trials with k consecutive successes
pass@k = (5) and reliability (pass@8 = 72.8%) through task-specific opti-
Total trials mization.
We evaluated pass@3, pass@5 and pass@8, with produc- Three agents form the Pareto frontier: ReAct-GPT-o3
tion deployment requiring pass@8 ≥ 80% for mission- (optimal cost), Plan-Execute (balanced), and Domain-Tuned
critical applications. Single-run evaluation masks brittle- (optimal reliability). Reflexion, despite the highest raw ac-
ness: an agent with 70% pass@1 might achieve only 30% curacy, is dominated by Plan-Execute, which achieves com-
pass@8, making it unsuitable for production despite accept- parable efficacy (71.9% vs 74.1%) at 4.1x lower cost with
able single-run performance. superior reliability (64.5% vs 61.2% pass@8).
Single-run success (pass@1) ranges from 68-74%, but
3.2 Composite Score and Enterprise Task Suite 8-run consistency (pass@8) drops dramatically to 52-73%.
For single-number comparison: CLEAR = wC · Cnorm + Domain-Tuned maintains 72.8% pass@8 (10.8% drop from
wL · Lnorm + wE · E + wA · A + wR · R where the weights pass@1), while ReAct-GPT4 drops 19.4% (from 72.3% to
sum to 1 and are application-specific. Cost and latency are 58.3%), revealing brittleness in general-purpose agents that
normalized via min-max scaling. Default equal weighting single-run evaluation masks.
(wi = 0.2) provides a balanced assessment, while enter- Domain-Tuned excels in compliance (+7.5 points over
prises can customize (e.g., financial services: wR = 0.4, ReAct-GPT4) and multi-stakeholder tasks (+7.5 points)
wA = 0.3; customer-facing: wL = 0.35). where specialized knowledge matters most. ReAct-GPT4
larger-scale studies would strengthen generalizability. Fu-
Table 2: Correlation with expert-rated deployment readiness ture work should explore adaptive weight learning, online
(N=15 experts, 40 tasks each). evaluation methodologies, multi-agent coordination evalua-
Evaluation Approach Pearson ρ Spearman ρs tion, interpretability metrics, and long-horizon evaluation.
Efficacy Only 0.41 0.39
Efficacy + Cost 0.58 0.56 6 Conclusion
CLEAR (All 5 Dimensions) 0.83 0.81 This work establishes multidimensional evaluation as es-
sential for enterprise agentic AI systems. Through system-
atic analysis of 12 main benchmarks, we identified critical
gaps: cost is unmeasured despite 50x variations, reliability
performs best in software development (73.3%). All agents
is untested despite drops in consistency 60 to 25%, and op-
met customer support SLAs (3 sec), but Plan-Execute and
erational requirements are absent despite 37% lab to pro-
Reflexion exceeded software development SLAs (30 sec) on
duction gaps. Our CLEAR framework addresses these lim-
23% and 34% of tasks due to extensive reflection. Domain-
itations with a comprehensive assessment of cost, latency,
Tuned achieved the highest adherence to policy (PAS = 0.93)
effectiveness, assurance, and reliability.
with the best resistance to prompt injection (8% successful
Empirical evaluation of six leading agents across 300
attacks vs. 18% for ToolFormer). See Appendix B for de-
enterprise tasks demonstrates that accuracy optimization
tailed domain breakdowns.
yields suboptimal deployments: agents with highest raw
Expert Validation. Table 2 shows CLEAR correlates
accuracy cost 4.4-10.8x more than Pareto-efficient alterna-
strongly with expert deployment readiness (ρ = 0.83,
tives. Domain-specialized approaches outperform general-
p<0.001), substantially outperforming efficacy-only evalu-
purpose architectures, achieving superior cost-normalized
ation (ρ = 0.41). The experts valued reliability the most,
performance (260.4 vs. 14.5-58.0 CNA) and reliability
followed by policy compliance and cost-efficiency. An ex-
(72.8% vs 52.1-64.5% pass@8). Expert validation confirms
pert noted: ”A 70% agent that works reliably is far more
that CLEAR predicts production success better (ρ = 0.83)
deployable than an 80% agent that is unpredictable and ex-
than traditional metrics (ρ = 0.41).
pensive.”
As agentic AI moves from research to production, eval-
uation must evolve accordingly. CLEAR provides a foun-
5 Discussion dation for enterprise-appropriate assessment, enabling cost-
Implications for Benchmark Design. Our findings demon- conscious architectural decisions and reliable deployment.
strate that academic benchmarks must evolve beyond the We plan to release our Enterprise Task Suite, evaluation
accuracy-centric evaluation. Cost transparency should be code, and complete experimental results to facilitate adop-
mandatory-reporting token consumption, API costs, and in- tion and enable reproducible research to advance practical
ference time enables fair comparison. Reliability assess- agent systems.
ment using pass @ k metrics (minimum k = 8) reveals
production-critical brittleness invisible in single-run evalu- References
ation. Domain-specific evaluation better captures real-world Beetz, R., and Riedl, Y. 2019. Robotic Process Automation:
performance than generic benchmarks, as evidenced by 15- Developing a Multi-Criteria Evaluation Model for the Se-
25% performance variations across domains. lection of Automatable Business Processes. In AMCIS 2019
Architectural Insights. Domain-tuned models demon- Proceedings, volume 4.
strate surprising effectiveness, achieving best cost- Chen, J.; Lu, Y.; Wang, X.; Zeng, H.; et al. 2024.
normalized accuracy and reliability despite using smaller Multi-Agent-as-Judge: Aligning LLM-Agent-Based Auto-
base models (70B vs GPT-4’s rumored 1.7T parameters). mated Evaluation with Multi-Dimensional Human Evalua-
Task-specific optimization outweighs the scale of the raw tion. arXiv preprint arXiv:2507.21028.
model for enterprise-focused applications. Conversely,
complex agent architectures (Reflexion with multi-turn Chiu, O. 2025. The Key to Production AI Agents: Evalua-
self-refinement) show diminishing returns. Although self- tions. Databricks Blog.
reflection improves accuracy marginally, it dramatically Deng, X.; Gu, Y.; Zheng, B.; Chen, S.; Stevens, S.; Wang,
increases cost and latency while harming reliability through B.; Sun, H.; and Su, Y. 2023. Mind2Web: Towards a Gen-
additional failure modes. The Plan-Execute agent’s strong eralist Agent for the Web. In NeurIPS Datasets and Bench-
performance (Pareto-optimal, balanced) indicates that marks Track.
hierarchical decomposition with cost-aware component Drouin, A.; Gasse, M.; Caccia, M.; Laradji, I. H.; Del
selection offers a practical middle ground. Verme, M.; Marty, T.; Boisvert, L.; Thakkar, M.; Cappart,
Limitations. Task coverage, while spanning six domains Q.; Vazquez, D.; et al. 2024. WorkArena: How Capable Are
with 300 tasks, remains limited compared to the real diver- Web Agents at Solving Common Knowledge Work Tasks?
sity of the company. Agent selection focused on established arXiv preprint arXiv:2403.07718.
architectures; newer approaches warrant investigation. The Jimenez, C. E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press,
cost dynamics will shift with the evolution of the model pric- O.; and Narasimhan, K. 2024. SWE-bench: Can Language
ing, though the relative comparisons remain valid. Human Models Resolve Real-World GitHub Issues? In Interna-
evaluation with 15 experts provides initial validation, but tional Conference on Learning Representations (ICLR).
Kang, D., et al. 2024. AI Agent Benchmarks Are Broken: Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan,
A Critical Analysis. Technical Report. K.; and Cao, Y. 2023. ReAct: Synergizing Reasoning and
Kapoor, S.; Stroebl, B.; Pandit, S.; Bommasani, R.; and Acting in Language Models. In International Conference on
Narayanan, A. 2024. AI Agents That Matter. arXiv preprint Learning Representations (ICLR).
arXiv:2407.01502. Yao, S.; Hu, J.; Shinn, N.; and Narasimhan, K. 2024. τ -
Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, bench: A Benchmark for Tool-Agent-User Interaction in
Y.; Ding, H.; Men, K.; Yang, K.; et al. 2023a. AgentBench: Real-World Domains. arXiv preprint arXiv:2406.12045.
Evaluating LLMs as Agents. In International Conference on Yehudai, A.; Eden, L.; Li, A.; Uziel, G.; Zhao, Y.; Bar-
Learning Representations (ICLR). Haim, R.; Cohan, A.; and Shmueli-Scheuer, M. 2025. Sur-
Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. vey on Evaluation of LLM-based Agents. arXiv preprint
2023b. G-eval: NLG Evaluation using GPT-4 with Better arXiv:2503.16416.
Human Alignment. In Conference on Empirical Methods in Zhang, Z., et al. 2024. Agent-SafetyBench: Evaluating the
Natural Language Processing (EMNLP). Safety of LLM Agents. arXiv preprint arXiv:2412.14470.
Liu, J.; Zhang, K.; Chen, W.; Alkhouli, T.; Negrinho, Zhou, S.; Xu, F. F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar,
R.; Elluru, V. R.; Benajiba, Y.; et al. 2024. Towards A.; Cheng, X.; Bisk, Y.; Fried, D.; Alon, U.; et al. 2023.
Effective GenAI Multi-Agent Collaboration: Design and WebArena: A Realistic Web Environment for Building Au-
Evaluation for Enterprise Applications. arXiv preprint tonomous Agents. In Advances in Neural Information Pro-
arXiv:2412.05449. cessing Systems, volume 36.
Mialon, G.; Fourrier, C.; Swift, C.; Wolf, T.; LeCun, Y.; and
Scialom, T. 2023. GAIA: A Benchmark for General AI A Enterprise Task Suite Details
Assistants. arXiv preprint arXiv:2311.12983.
Our benchmark comprises 300 tasks across six enterprise
Patil, S. G.; Yan, F.; Mao, H.; Ji, C. C.-W.; Zhang, T.; Sto- domains, each with ground-truth annotations for all CLEAR
ica, I.; and Gonzalez, J. E. 2024. Gorilla: Large Language dimensions:
Model Connected with Massive APIs. Berkeley Function Customer Support (60 tasks): Multi-turn policy-
Calling Leaderboard. compliant issue resolution requiring knowledge base re-
Qin, Y.; Liang, S.; Ye, Y.; Zhu, K.; Yan, L.; Lu, Y.; Lin, Y.; trieval with proper citation, escalation handling with role-
Cong, X.; Tang, X.; Qian, B.; et al. 2024. ToolLLM: Fa- based routing, and complaint management with sentiment
cilitating Large Language Models to Master 16000+ Real- analysis. Tasks span routine inquiries (30%), complex tech-
world APIs. In International Conference on Learning Rep- nical issues (40%), and edge cases requiring human escala-
resentations (ICLR). tion (30%). SLA: 3 seconds average response time.
Aisera Research Team. 2024. CLASSic: A Holistic Frame- Data Analysis (50 tasks): Report generation from struc-
work for Evaluating Enterprise AI Agents. Aisera Technical tured databases, SQL query construction with optimization
Report. constraints, data visualization with accuracy verification,
and trend analysis with statistical validation. Tasks require
Salesforce AI Research. 2024. First-Of-Its-Kind LLM handling missing data, outlier detection, and ensuring com-
Benchmark Ranks Generative AI Against Real-World Busi- pliance with data privacy regulations. SLA: 15 seconds for
ness Tasks. Salesforce Blog. query execution, 45 seconds for reports.
Samsung Research. 2024. TRUEBench: Trustworthy Real- Process Automation (50 tasks): Form completion with
world Usage Evaluation Benchmark for Enterprise LLMs. validation rules, approval workflow navigation across multi-
Samsung AI Research. step processes, cross-system integration requiring API or-
Schick, T.; Dwivedi-Yu, J.; Dessı̀, R.; Raileanu, R.; Lomeli, chestration, and exception handling for edge cases. Tasks
M.; Zettlemoyer, L.; Cancedda, N.; and Scialom, T. 2023. test error recovery, rollback mechanisms, and audit trail gen-
Toolformer: Language Models Can Teach Themselves to eration. SLA: 10 seconds per process step.
Use Tools. In NeurIPS. Software Development (60 tasks): Bug fixing in pro-
Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; and duction codebases with test-driven validation, code review
Yao, S. 2023. Reflexion: Language Agents with Verbal with security vulnerability detection, test generation achiev-
Reinforcement Learning. In NeurIPS. ing ¿80% coverage, and refactoring with performance opti-
mization. Tasks span Python (40%), JavaScript (30%), and
Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Hua, Java (30%) from real enterprise repositories. SLA: 30 sec-
T. J.; Shin, D.; Lei, F.; Liu, Y.; et al. 2024. OSWorld: Bench- onds for analysis, 60 seconds for code generation.
marking Multimodal Agents for Open-Ended Tasks in Real Compliance (40 tasks): GDPR request processing (data
Computer Environments. In NeurIPS Datasets and Bench- access, deletion, portability), audit trail validation ensuring
marks Track. complete traceability, regulatory requirement verification
Yang, J.; Prabhakar, A.; Narasimhan, K.; and Yao, S. against SOC 2 and ISO 27001 standards, and policy enforce-
2023. InterCode: Standardizing and Benchmarking Interac- ment across multi-stakeholder workflows. Tasks require pre-
tive Coding with Execution Feedback. In NeurIPS Datasets cise interpretation of legal language and zero-tolerance for
and Benchmarks Track. policy violations. SLA: 20 seconds per compliance check.
Multi-Stakeholder Workflows (40 tasks): Cross- ing, particularly for complex refactoring. However, policy
departmental coordination requiring approval chains, con- violations occur through unsafe code patterns (hardcoded
flict resolution with competing stakeholder priorities, role- credentials, SQL injection vulnerabilities). Domain-Tuned
based access control enforcement, and deadline manage- balances code quality with security compliance (PAS 0.94).
ment with escalation paths. These represent the most com- Compliance: Largest performance gap across agents (65-
plex scenarios with 8-15 steps, multiple decision points, 72.5%). Domain-Tuned significantly outperforms (+7.5
and requiring negotiation between conflicting requirements. points) due to training on regulatory documents and legal
SLA: 15 seconds per interaction. language. General-purpose agents struggle with precise le-
All tasks include: (1) Natural language task description, gal interpretation, often providing approximately correct but
(2) Required input data/context, (3) Expected output with legally insufficient responses. PAS variation (0.82-0.93) re-
ground-truth, (4) Policy documents defining constraints, (5) flects challenges in zero-tolerance compliance requirements.
Cost baselines from reference implementations, (6) Latency Multi-Stakeholder: Lowest efficacy across all agents (61-
SLA thresholds, (7) Security test cases, (8) Reliability base- 69%), highlighting fundamental challenges in complex co-
line from 50 human expert executions. ordination. Tasks require balancing conflicting priorities,
managing approval chains with 3-5 stakeholders, and ne-
B Detailed Domain Performance gotiating constraint violations. Domain-Tuned’s advantage
(+7.5 points over ReAct-GPT4) comes from learned conflict
Table 3 presents a comprehensive breakdown of the domain-
resolution patterns. Low PAS (0.78-0.89) indicates frequent
specific performance in the six agents and the dimensions of
policy violations in complex scenarios with competing re-
efficacy and policy adherence.
quirements.

Table 3: Detailed domain-specific performance showing ef- C Cost and Latency Analysis
ficacy (%) and policy adherence score (PAS) across six en-
terprise task categories.
ReAct-GPT4 Plan-Exec Domain-T Table 4: Detailed cost breakdown and latency analysis
Domain Eff. PAS Eff. PAS Eff. PAS across agents.
Customer Support 78.3 0.87 75.0 0.85 81.7 0.95 Agent Input Output Total Plan Exec Reflect
Data Analysis 69.0 0.94 72.0 0.93 71.0 0.96 Tok. Tok. Cost (s) (s) (s)
Process Automation 71.0 0.88 73.0 0.89 72.0 0.92 ReAct-GPT4 47.2K 8.3K $2.87 2.1 4.8 1.5
Software Dev. 73.3 0.91 70.0 0.87 71.7 0.94 ReAct-GPT-o3 52.1K 9.7K $0.31 1.2 2.4 0.6
Compliance 65.0 0.82 67.5 0.84 72.5 0.93 Reflexion 89.4K 15.2K $5.12 3.4 6.1 3.2
Multi-Stakeholder 61.3 0.78 64.0 0.81 68.8 0.89 Plan-Execute 38.6K 7.1K $1.24 1.8 4.2 0.8
Overall 72.3 0.89 71.9 0.88 70.3 0.93 ToolFormer 44.3K 9.8K $1.89 1.5 3.6 0.8
Domain-Tuned 31.2K 5.4K $0.27 0.9 2.3 0.6

Key Domain Insights:


Table 4 reveals that Reflexion’s high cost stems from mul-
Customer Support: Domain-Tuned excels due to fine-
tiple self-reflection iterations (avg 2.8 iterations per task),
tuning on enterprise support tickets with policy-compliant
increasing both token consumption and latency. Domain-
response templates. The 81.7% efficacy represents success-
Tuned achieves lowest cost through efficient inference on
ful resolution without escalation. PAS of 0.95 indicates min-
smaller model with task-specific optimization reducing un-
imal policy violations (e.g., unauthorized discounts, data
necessary reasoning steps. Plan-Execute balances cost by
sharing). ReAct-GPT4 struggles with policy boundaries
using GPT-4 for planning (15% of tokens) and GPT-o3 for
(PAS 0.87), often over-promising or providing unauthorized
execution (85% of tokens).
information.
Latency breakdown shows reflection phases dominate for
Data Analysis: All agents perform moderately well (69- Reflexion (3.2s, 25% of total), while Domain-Tuned mini-
72% efficacy), with high policy adherence (0.93-0.96) since mizes reflection through learned patterns. Planning latency
SQL constraints naturally enforce data access policies. Plan- varies 0.9-3.4s, with Reflexion requiring complex multi-step
Execute slightly outperforms (72%) through hierarchical planning. These findings inform architectural decisions: re-
query decomposition. Failures primarily stem from complex flection loops provide marginal accuracy gains at dispropor-
multi-table joins and aggregate functions. tionate cost and latency penalties.
Process Automation: Balanced performance across
agents (71-73%). Plan-Execute’s hierarchical structure
aligns naturally with multi-step workflows. Policy adher-
References
ence challenges include skipping required approvals and Beetz, R., and Riedl, Y. 2019. Robotic Process Automation:
violating temporal sequencing constraints. Domain-Tuned Developing a Multi-Criteria Evaluation Model for the Se-
achieves highest PAS (0.92) through learned workflow pat- lection of Automatable Business Processes. In AMCIS 2019
terns. Proceedings, volume 4.
Software Development: ReAct-GPT4 achieves best effi- Chen, J.; Lu, Y.; Wang, X.; Zeng, H.; et al. 2024.
cacy (73.3%) leveraging GPT-4’s superior code understand- Multi-Agent-as-Judge: Aligning LLM-Agent-Based Auto-
mated Evaluation with Multi-Dimensional Human Evalua- Schick, T.; Dwivedi-Yu, J.; Dessı̀, R.; Raileanu, R.; Lomeli,
tion. arXiv preprint arXiv:2507.21028. M.; Zettlemoyer, L.; Cancedda, N.; and Scialom, T. 2023.
Chiu, O. 2025. The Key to Production AI Agents: Evalua- Toolformer: Language Models Can Teach Themselves to
tions. Databricks Blog. Use Tools. In NeurIPS.
Deng, X.; Gu, Y.; Zheng, B.; Chen, S.; Stevens, S.; Wang, Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; and
B.; Sun, H.; and Su, Y. 2023. Mind2Web: Towards a Gen- Yao, S. 2023. Reflexion: Language Agents with Verbal
eralist Agent for the Web. In NeurIPS Datasets and Bench- Reinforcement Learning. In NeurIPS.
marks Track. Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Hua,
Drouin, A.; Gasse, M.; Caccia, M.; Laradji, I. H.; Del T. J.; Shin, D.; Lei, F.; Liu, Y.; et al. 2024. OSWorld: Bench-
Verme, M.; Marty, T.; Boisvert, L.; Thakkar, M.; Cappart, marking Multimodal Agents for Open-Ended Tasks in Real
Q.; Vazquez, D.; et al. 2024. WorkArena: How Capable Are Computer Environments. In NeurIPS Datasets and Bench-
Web Agents at Solving Common Knowledge Work Tasks? marks Track.
arXiv preprint arXiv:2403.07718. Yang, J.; Prabhakar, A.; Narasimhan, K.; and Yao, S.
Jimenez, C. E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, 2023. InterCode: Standardizing and Benchmarking Interac-
O.; and Narasimhan, K. 2024. SWE-bench: Can Language tive Coding with Execution Feedback. In NeurIPS Datasets
Models Resolve Real-World GitHub Issues? In Interna- and Benchmarks Track.
tional Conference on Learning Representations (ICLR). Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan,
Kang, D., et al. 2024. AI Agent Benchmarks Are Broken: K.; and Cao, Y. 2023. ReAct: Synergizing Reasoning and
A Critical Analysis. Technical Report. Acting in Language Models. In International Conference on
Kapoor, S.; Stroebl, B.; Pandit, S.; Bommasani, R.; and Learning Representations (ICLR).
Narayanan, A. 2024. AI Agents That Matter. arXiv preprint Yao, S.; Hu, J.; Shinn, N.; and Narasimhan, K. 2024. τ -
arXiv:2407.01502. bench: A Benchmark for Tool-Agent-User Interaction in
Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, Real-World Domains. arXiv preprint arXiv:2406.12045.
Y.; Ding, H.; Men, K.; Yang, K.; et al. 2023a. AgentBench: Yehudai, A.; Eden, L.; Li, A.; Uziel, G.; Zhao, Y.; Bar-
Evaluating LLMs as Agents. In International Conference on Haim, R.; Cohan, A.; and Shmueli-Scheuer, M. 2025. Sur-
Learning Representations (ICLR). vey on Evaluation of LLM-based Agents. arXiv preprint
Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. arXiv:2503.16416.
2023b. G-eval: NLG Evaluation using GPT-4 with Better Zhang, Z., et al. 2024. Agent-SafetyBench: Evaluating the
Human Alignment. In Conference on Empirical Methods in Safety of LLM Agents. arXiv preprint arXiv:2412.14470.
Natural Language Processing (EMNLP). Zhou, S.; Xu, F. F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar,
Liu, J.; Zhang, K.; Chen, W.; Alkhouli, T.; Negrinho, A.; Cheng, X.; Bisk, Y.; Fried, D.; Alon, U.; et al. 2023.
R.; Elluru, V. R.; Benajiba, Y.; et al. 2024. Towards WebArena: A Realistic Web Environment for Building Au-
Effective GenAI Multi-Agent Collaboration: Design and tonomous Agents. In Advances in Neural Information Pro-
Evaluation for Enterprise Applications. arXiv preprint cessing Systems, volume 36.
arXiv:2412.05449.
Mialon, G.; Fourrier, C.; Swift, C.; Wolf, T.; LeCun, Y.; and
Scialom, T. 2023. GAIA: A Benchmark for General AI
Assistants. arXiv preprint arXiv:2311.12983.
Patil, S. G.; Yan, F.; Mao, H.; Ji, C. C.-W.; Zhang, T.; Sto-
ica, I.; and Gonzalez, J. E. 2024. Gorilla: Large Language
Model Connected with Massive APIs. Berkeley Function
Calling Leaderboard.
Qin, Y.; Liang, S.; Ye, Y.; Zhu, K.; Yan, L.; Lu, Y.; Lin, Y.;
Cong, X.; Tang, X.; Qian, B.; et al. 2024. ToolLLM: Fa-
cilitating Large Language Models to Master 16000+ Real-
world APIs. In International Conference on Learning Rep-
resentations (ICLR).
Aisera Research Team. 2024. CLASSic: A Holistic Frame-
work for Evaluating Enterprise AI Agents. Aisera Technical
Report.
Salesforce AI Research. 2024. First-Of-Its-Kind LLM
Benchmark Ranks Generative AI Against Real-World Busi-
ness Tasks. Salesforce Blog.
Samsung Research. 2024. TRUEBench: Trustworthy Real-
world Usage Evaluation Benchmark for Enterprise LLMs.
Samsung AI Research.

You might also like