Benchmark-Driven Evaluation Frameworks for AI Agent Trustworthiness in Enterprise Analytics

Authors

  • Santosh Reddy Muthyala

Keywords:

AI agent trustworthiness; enterprise analytics; text-to-SQL evaluation; benchmark design; LLM-as-judge; responsible AI adoption

Abstract

Enterprise adoption of AI-powered analytics agents has outpaced the methods available to evaluate whether their outputs can be trusted. General-purpose benchmarks such as MMLU [1] and HumanEval [2] measure static knowledge recall and isolated code-generation ability, while text-to-SQL benchmarks such as Spider [3] and BIRD [4] test query generation against simplified schemas. None of these evaluate the full chain of reasoning an enterprise analytics agent must perform: schema interpretation, organization-specific metric computation, calibrated uncertainty, and application of the correct analytical methodology. On BIRD, even the strongest reported system reaches only 40.08% execution accuracy against realistic enterprise schemas, compared with 92.96% for human analysts [4], a gap that becomes consequential once an agent's output informs financial or regulatory decisions rather than exploratory analysis. This article proposes a nine-dimension trustworthiness evaluation framework purpose-built for enterprise analytics agents. Four dimensions are grounded in established trust taxonomies [6]-[8]; five are analytics-specific contributions not addressed by any existing framework: Schema Fidelity, Metric Definition Adherence, Failure Mode Transparency, Business Context Understanding, and Analytical Methodology Competence. The article also proposes a benchmark design methodology combining automated, LLM-as-judge, and domain-expert graders, and a five-phase Continuous Benchmark Evolution Lifecycle for keeping evaluation infrastructure current as schemas, metric definitions, and agent capabilities change.

 

Downloads

Download data is not yet available.

References

D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, "Measuring Massive Multitask Language Understanding," in Proc. ICLR, 2021. arXiv:2009.03300.

M. Chen, J. Tworek, H. Jun, et al., "Evaluating Large Language Models Trained on Code," arXiv:2107.03374, 2021.

T. Yu, R. Zhang, K. Yang, et al., "Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task," in Proc. EMNLP, 2018. DOI: 10.18653/v1/D18-1425.

J. Li, B. Hui, G. Qu, et al., "Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQL Evaluation," in Proc. NeurIPS Datasets and Benchmarks Track, 2023. arXiv:2305.03111.

P. Liang, R. Bommasani, et al., "Holistic Evaluation of Language Models," arXiv:2211.09110, 2022.

National Institute of Standards and Technology, "AI Risk Management Framework (AI RMF 1.0)," NIST AI 100-1, 2023. DOI: 10.6028/NIST.AI.100-1.

B. Wang, W. Chen, H. Pei, et al., "DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models," in Proc. NeurIPS, 2023. arXiv:2306.11698.

Y. Huang, L. Sun, H. Wang, et al., "TrustLLM: Trustworthiness in Large Language Models," arXiv:2401.05561, 2024.

Accenture, "From Compliance to Confidence: Embracing a New Mindset to Advance Responsible AI Maturity," Accenture Research, 2024.

Accenture, "Technology Vision 2025," Accenture Research, 2025.

L. Zheng, W. Chiang, Y. Sheng, et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena," in Proc. NeurIPS Datasets and Benchmarks Track, 2023. arXiv:2306.05685.

X. Liu, H. Yu, H. Zhang, et al., "AgentBench: Evaluating LLMs as Agents," in Proc. ICLR, 2024. arXiv:2308.03688.

S. Chang, J. Wang, M. Dong, et al., "Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness," in Proc. ICLR, 2023. arXiv:2301.08881.

D. Gao, H. Wang, Y. Li, et al., "Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation," arXiv:2308.15363, 2023.

X. Wang, Z. Wang, J. Liu, et al., "MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback," in Proc. ICLR, 2024. arXiv:2309.10691.

M. Pourreza and D. Rafiei, "DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction," in Proc. NeurIPS, 2023. arXiv:2304.11015.

Y. Lai, C. Li, Y. Wang, et al., "DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation," in Proc. ICML, 2023. arXiv:2211.11501.

L. Huang, W. Yu, W. Ma, et al., "A Survey on Hallucination in Large Language Models," arXiv:2311.05232, 2023.

Downloads

Published

31.08.2026

How to Cite

Santosh Reddy Muthyala. (2026). Benchmark-Driven Evaluation Frameworks for AI Agent Trustworthiness in Enterprise Analytics. International Journal of Intelligent Systems and Applications in Engineering, 14(1s), 2315–2324. Retrieved from https://www.ijisae.org/index.php/IJISAE/article/view/8531

Issue

Section

Research Article