Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Finauditing: a financial taxonomy-structured multi-document benchmark for evaluating LLMs

Wang, Yan, Wang, Keyi, Yang, Shanshan, Patel, Jaisal, Zhao, Jeff, Mo, Fengran, Peng, Xueqing, Qian, Lingfei, Chen, Yankai, Gutiérrez-Basulto, Víctor ORCID: https://orcid.org/0000-0002-6117-5459, Huang, Jimin, Xiong, Guojun, Liu, Xiao-Yang and Nie, Jian-Yun 2026. Finauditing: a financial taxonomy-structured multi-document benchmark for evaluating LLMs. Presented at: SIGIR '26: The 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, Melbourne, Australia, 20-24 July 2026. SIGIR '26: Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval. New York, NY: ACM, pp. 3456-3463. 10.1145/3805712.3808578

[thumbnail of 3805712.3808578.pdf] PDF - Published Version
Available under License Creative Commons Attribution.

Download (9MB)

Abstract

Going beyond simple text processing, financial auditing requires detecting semantic, structural, and numerical inconsistencies across large-scale disclosures. As financial reports are filed in XBRL, a structured XML format governed by accounting standards, auditing becomes a structured information extraction and reasoning problem involving concept alignment, taxonomy-defined relations, and cross-document consistency. Although large language models (LLMs) show promise on isolated financial tasks, their capability in professional-grade auditing remains unclear. We introduce FinAuditing, a taxonomy-aligned, structure-aware benchmark built from real XBRL filings. It contains 1,102 annotated instances averaging over 33k tokens and defines three tasks: Financial Semantic Matching (FinSM), Financial Relationship Extraction (FinRE), and Financial Mathematical Reasoning (FinMR). Evaluations of 13 state-of-the-art LLMs reveal substantial gaps in concept retrieval, taxonomy-aware relation modeling, and consistent cross-document reasoning. These findings highlight the need for realistic, structure-aware benchmarks. We release the evaluation code1 and dataset2 publicly, and the task currently serves as the official benchmark of an ongoing public evaluation contest3.

Item Type: Conference or Workshop Item - published (Paper)
Date Type: Published Online
Status: Published
Schools: Schools > Computer Science & Informatics
Publisher: ACM
ISBN: 9798400725999
Date of First Compliant Deposit: 22 July 2026
Last Modified: 22 Jul 2026 13:30
URI: https://orca.cardiff.ac.uk/id/eprint/188425

Actions (repository staff only)

Edit Item Edit Item

Downloads

Downloads per month over past year

View more statistics