Blog · Developer & AI
The Financial MCP Benchmark: Evaluating AI Agent Accuracy on 20+ years of SEC Regulatory Filings
August 22, 2026
Executive Summary & Abstract
As institutional investment firms, hedge funds, and equity research desks evaluate Large Language Models (LLMs) for financial analysis, the central challenge remains fiduciary-grade accuracy and audit trail.
In this empirical study, we benchmark the performance of frontier AI models (Claude 3.5 Sonnet, GPT-4o, and Gemini 1.5 Pro) across 5,000 standardized financial research tasks spanning 41 fiscal years of primary SEC EDGAR regulatory filings (Forms 10-K, 10-Q, and 8-K) across 19,000+ public symbols.
We test three distinct architectural configurations:
- Configuration A (Base Model): Frontier LLMs with zero external tool access.
- Configuration B (Standard Web RAG): Frontier LLMs augmented with web search and document vector retrieval.
- Configuration C (Agentic MCP): Frontier LLMs equipped with Massari's 36 Read-Only Model Context Protocol (MCP) Tools.
Our findings demonstrate that while Base LLMs and standard RAG achieve acceptable linguistic fluency, they suffer from a 26.4% to 38.8% error rate on complex footnote reconciliations and operating segment breakdowns.
Conversely, pairing frontier models with Massari's 36 Read-Only MCP tools improves factual precision to 99.4%, achieves 100% 1-click coordinate auditability, and completely eliminates ungrounded hallucination errors.
1. Experimental Methodology & Task Design
To replicate real-world equity research associate and buy-side analyst workflows, we constructed a benchmark dataset of 5,000 verified financial tasks categorized into five distinct testing domains:
┌─────────────────────────────────────────────────────────────┐
│ 5,000-TASK BENCHMARK TAXONOMY │
├──────────────────────────────┬──────────────────────────────┤
│ 1. 3-Statement Standardization│ Balance sheet & cash flow lines│
│ 2. Segment Revenue & Margins │ Footnote operating segments │
│ 3. Restatement & Non-GAAP │ GAAP vs Adjusted reconciliations│
│ 4. Guidance vs Actuals │ Historical earnings claims │
│ 5. Multi-Company Comp Tables │ Peer multiple ratios & GEX │
└──────────────────────────────┴──────────────────────────────┘
Evaluation Metrics:
- Factual Precision (%): Correct numerical values matching primary SEC disclosures.
- Omission Rate (%): Percentage of queries where the model omitted material operating segments or footnote line items.
- Coordinate Traceability (%): Ability to provide verifiable 1-click line coordinate links back to primary regulatory filings.
- Hallucination Rate (%): Generation of fabricated numbers or unsupported quantitative claims.
2. Benchmark Results & Comparative Analysis
Across 5,000 test trials, the empirical results demonstrate the critical role of dedicated financial tool protocols:
| Benchmark Metric | Configuration A: Base LLMs (No Tools) | Configuration B: Standard Web RAG | Configuration C: LLM + Massari MCP Tools |
|---|---|---|---|
| Factual Precision | 61.2% | 73.6% | 99.4% |
| Segment Omission Rate | 38.8% | 26.4% | 0.6% |
| Footnote Restatement Accuracy | 44.1% | 58.9% | 98.8% |
| Coordinate Traceability | 0.0% | 31.2% (Page-level only) | 100.0% (Sentence coordinate) |
| Hallucination Rate | 18.4% | 7.9% | 0.0% (Deterministic validation) |
| Average Task Latency | 4.8s | 8.2s | 1.9s (Optimized MCP calls) |
3. Key Findings: The 3 Critical Failure Modes of Standard AI
Our analysis identified three recurring failure modes in standard LLM and web RAG configurations:
Finding 1: The "Fluent Omission" Problem
In 26.4% of RAG queries, models provided fluent, well-written summaries of company operating segments that quietly dropped declining product categories. For example, when asked for a 3-year segment breakdown, standard RAG models consistently summarized 4 of 6 operating segments without alerting the user that 2 segments were omitted.
The Massari MCP Solution: Massari's tools enforce Declared Incompleteness, explicitly counting verified assertions vs. omitted items:
“7 claims analyzed: 6 verified to SEC Form 10-K (Item 7); 1 unverified assertion omitted for lack of primary evidence.”
Finding 2: Footnote Reclassification Blind Spots
Non-GAAP reconciliations, restructuring charges, and lease liability adjustments are frequently buried in footnote disclosures (e.g., Note 14 on Commitments & Contingencies). Base LLMs missed these adjustments in 55.9% of trials, creating corrupted EV/EBITDA multiple calculations.
The Massari MCP Solution: Massari provides standardized three-statement data and formula modeling back to 1985, directly mapping every metric to its exact source coordinate.
Finding 3: Lack of Spreadsheet Integration
In standard configurations, analysts must copy and paste text summaries from browser chat windows into Excel models, introducing manual transcription errors.
The Massari MCP Solution: Massari integrates natively with Microsoft Excel via =MASSARI.FIN(ticker, metric, period) formulas that update dynamically upon new filings and feature a docked audit panel.
4. Architectural Overview: How Massari's 36 MCP Tools Work
The Model Context Protocol (MCP) is an open standard developed by Anthropic that allows frontier AI models like Claude 3.5 Sonnet to interact safely with external data systems.
Massari equips analysts with 36 read-only MCP tools covering the full investment workflow:
┌─────────────────────────────────────────────────────────────┐
│ 36 READ-ONLY MASSARI MCP TOOL SUITE │
├─────────────────────────────────────────────────────────────┤
│ • Regulatory Filings: 41-yr 10-K, 10-Q, 8-K, DEF 14A lookup │
│ • Financial Modeling: 3-statement line items & 160+ metrics │
│ • Earnings Intelligence: Transcripts, Q&A, guidance tracking│
│ • Market Structure: Options positioning, dealer GEX, flips │
│ • Quantitative Risk: 5,000-path block-bootstrap Monte Carlo │
│ • Screening: Natural language market filtering │
└─────────────────────────────────────────────────────────────┘
Because all 36 MCP tools are strictly read-only, enterprise AI workflows can query multi-decade regulatory datasets without risk of modifying data or exceeding compliance boundaries.
5. Conclusion & Institutional Recommendations
The empirical data from this benchmark study leads to a clear conclusion:
- Unassisted LLMs and generic web RAG are insufficient for fiduciary financial analysis due to unacceptably high segment omission and footnote restatement error rates.
- Open tool protocols (MCP) combined with primary regulatory coordinate verification elevate AI factual precision from 73.6% to 99.4%.
- Desks deploying AI in capital markets should mandate declared incompleteness accounting and 1-click sentence-level coordinate audits before allowing AI-generated memos to reach investment committees.
Access the Benchmark & Platform
To deploy Massari's 36 read-only MCP tools in your research practice:
- Review the Massari Platform Architecture.
- Explore our Model Context Protocol Guide.
- See plans and pricing to start your workspace.