A Methodology for Assessing the Risk of Metric Failure in LLMs Within the Financial Domain

As Generative Artificial Intelligence is adopted across the financial services industry, a significant barrier to adoption and usage is measuring model performance. Historical machine learning metrics can oftentimes fail to generalize to GenAI workloads and are often supplemented using Subject Matter Expert (SME) Evaluation. Even in this combination, many projects fail to account for various unique risks present in choosing specific metrics. Additionally, many widespread benchmarks created by foundational research labs and educational institutions fail to generalize to industrial use. This paper explains these challenges and provides a Risk Assessment Framework to allow for better application of SME and machine learning Metrics

Paper

References (13)

02Class action alleges unitedhealth used ai to wrongly deny medicare advantage claims2025 · Best’s News
03Making generative ai trustworthy and reliable for adoption at scale2025 · Turing Institute Blog
04Cigna sued over alleged automated patient claims denials2023 · Bloomberg Law
05Report on apple card investigation2021
07U.S. Congress. Equal Credit Opportunity Act (ECOA). 15 U.S.C. § 1691,1974 · tinyurl.com/USECOA
08Haven Health ManagementHaven Health Management Blog
09’Digital Workers’ Have Arrived in BankingThe Wall Street Journal
10High-level summary of the ai actEU Artificial Intelligence Act Website
11Building ai trust: The key role of explainabilityMcKinsey & Company Insights
12BNY, America’s Oldest Bank, Signs Multiyear Deal With OpenAIThe Wall Street Journal

Scroll for more · 1 remaining

Similar papers

© 2026 NYSGPT2525 LLC