One terminal. Financials, earnings, portfolios, reports, news, and more. Every number source-linked.

Blog · Developer & AI

The Wall Street 1,000: Benchmarking Frontier AI on Financial Retrieval and Falsification

August 22, 2026

EMPIRICAL BENCHMARKMassari Research Desk · 12 min read · 100 S&P 500 / Nasdaq 100 Equities · Dataset v2.0-US

Executive Summary & Abstract

Financial data is unforgiving. A misplaced decimal point, an unadjusted segment recast, or an inverted legal interpretation can cascade into a material misstatement.

In 2024, financial data provider Daloopa published a pioneering study evaluating Large Language Models on financial retrieval. Testing 500 quantitative questions across corporate filings, Daloopa demonstrated that out-of-the-box LLMs suffered from frequent rounding drift, calendar/fiscal period confusion, and table hallucination—concluding that general-purpose AI could not be trusted for institutional modeling without proprietary extraction layers.

Two years later, the AI landscape has transformed. Frontier reasoning models (such as ChatGPT 5.5, Gemini 3.7, and Claude Opus) paired with open standards like Anthropic's Model Context Protocol (MCP) have redefined what is possible in automated equity research.

To measure where financial AI stands today, we engineered The Wall Street 1,000—an open empirical benchmark expanding upon Daloopa’s foundation. Evaluating 1,000 fundamental research tasks across 100 S&P 500 and Nasdaq 100 companies, our study tests models across six core institutional pillars: Segment Recasts, Geographic Revenue, Sovereign Trade Risks, Supply Chain Commitments, Complex Legal Contingencies, and Sell-Side Earnings Call Q&A.

1,000
Evaluated Tasks
100 S&P 500 & Nasdaq 100 US domestic equities
0.0%
EDGAR Transcript Access
Unassisted LLMs fail 100% of earnings call Q&A
99.9%
10-K Table Precision
Reading backward-looking tables is now solved
~1.2s
MCP Retrieval Latency
Compared to 22-minute unassisted web crawls

Overall Benchmark Accuracy Across 1,000 Tasks (Exact Match 0% Tolerance)

Massari MCP Native Layer
100.0%
Gemini 3.7 + Massari MCP
100.0%
ChatGPT 5.5 + Massari MCP
85.7%
Gemini 3.7 (Unassisted Web)
80.0%
ChatGPT 5.5 (Unassisted Codex)
79.9%
Parametric LLM (Raw Memory)
32.0%
Claude Opus (Unassisted UI)
8.0%

1. Background: Daloopa's 500-Question Study and the Evolving AI Landscape

What Daloopa’s Benchmark Established

Daloopa’s original financial retrieval benchmark was an important contribution to the AI finance literature. By testing leading models on 500 discrete quantitative questions (such as reported Adjusted EBITDA or segment revenues), their research highlighted three critical vulnerabilities of early LLMs:

  1. The Rounding & Interpretation Gap: Generic models rounded off material decimals or confused reported metrics with non-GAAP adjustments.
  2. Fiscal vs. Calendar Period Confusion: Companies with non-standard fiscal years (e.g. NVIDIA or Apple) frequently caused models to extract numbers shifted by one or two quarters.
  3. Table Extraction Failures: Dense multi-column tables in unstructured PDFs often led models to drop rows or hallucinate values.

Daloopa concluded that general-purpose AI lacked the precision necessary for financial workflows and that dedicated tabular pipelines were essential.

What Has Changed in 2026?

Our evaluation reveals a major shift in frontier AI capabilities:

  • Table Extraction is Solved: In our 1,000-task evaluation, modern frontier models (ChatGPT 5.5 in Codex and Gemini 3.7) extracted 10-K and 10-Q filing figures with 99.9% to 100% accuracy, overcoming the table-reading limitations identified in 2024.
  • The New Bottleneck is Transcripts & Context Protocols: The primary barrier in 2026 is no longer parsing 10-K tables—it is access to real-time speaker-diarized earnings call transcripts, sub-second retrieval latency, and open Model Context Protocol (MCP) integrations.
“In capital markets, an AI model that invents a single SEC accession number is not an inaccurate assistant—it is a catastrophic compliance violation under SEC Rule 10b-5.”
— Massari Quantitative Research Team

2. The 3 Tiers of Financial AI Execution

Tier 1
Conversational Chat UI
Pure parametric memory. 8%–32% score with high hallucination risk on legal & accession codes.
Tier 2
Unassisted Coding Agent
Python web crawler. 79.9% score on SEC filings, but stalls for 22m and misses 100% of audio transcripts.
Tier 3
Massari MCP Terminal Layer
Grounded Model Context Protocol. 100% ground truth in ~1.2s across 20+yr filings & 19,000+ transcripts.

Key Takeaways

1
The 20% Transcript Hard Ceiling
Every unassisted AI hits a hard 80% ceiling. SEC EDGAR contains zero earnings call audio or Q&A transcripts. Unassisted models fail 100% of forward guidance and analyst exchanges.
2
Beyond Table Extraction
While 2024 benchmarks focused on 10-K tables, modern research requires multi-statement recasts, legal contingencies, and verbatim CFO guidance ranges.
3
Accession Numbers as Lie Detectors
High-entropy SEC Accession Numbers (CIK-YY-Sequence) cannot be guessed from parametric memory. Verifying cited accessions against EDGAR instantly detects synthetic hallucinations.
4
The Latency Chasm (22m vs 1.2s)
Unassisted web scraping takes 10 to 22 minutes and hits strict 10 req/sec EDGAR rate limits. Massari MCP delivers verified primary data in ~1.2 seconds per tool call.

3. Benchmark Architecture: The 6 Research Pillars

While early financial benchmarks focused predominantly on single-line income statement metrics, institutional equity research requires evaluating qualitative footnotes, forward guidance, and capital allocation disclosures.

The Wall Street 1,000 tests across six distinct pillars:

Research Pillar Task Count Evaluation Focus Accounting Standard
1. Revenue Segments & Recasts 200 Tasks Segment operating income & boundary recasts ASC 280 / ASU 2023-07
2. Geographic Revenue & FX Drag 200 Tasks Destination vs billed; Constant-currency FX ASC 830 / Non-GAAP MD&A
3. Sovereign & Export Risks 100 Tasks BIS export controls, tariffs, entity lists Item 1A / Export Admin
4. Supply Chain & Commitments 100 Tasks Unconditional purchase debt & capacity ASC 440 Commitments
5. Qualitative Legal & Tax Risks 200 Tasks Material litigation & IRS statutory notices ASC 450 Contingencies
6. Analyst-Led Earnings Call Q&A 200 Tasks Verbatim CFO guidance & sell-side exchanges Diarized Transcripts

4. Master Institutional Leaderboard

Across all 1,000 tasks, we benchmarked models on exact match precision, tolerance bands, transcript coverage, and institutional compliance trust:

Model / Configuration Exact Match (0% Tol) 1% Tolerance Transcripts (Q&A) Institutional Trust Score
Massari MCP Native Layer 100.0% 100.0% 100.0% Verified 100.0% Zero Risk
Gemini 3.7 + Massari MCP 100.0% 100.0% 100.0% Verified 100.0% Zero Risk
ChatGPT 5.5 + Massari MCP 85.7% 85.7% 85.0% Grounded 85.7% Grounded
Gemini 3.7 (Unassisted Web) 80.0% 84.5% 0.0% Failed 80.0% Missing Q&A
ChatGPT 5.5 (Unassisted Codex) 79.9% 81.2% 0.0% Failed 79.9% Missing Q&A
Claude Opus (Unassisted UI) 8.0% 8.0% 0.0% Timed Out 6.4% Incomplete
Parametric LLM (Raw Memory) 32.0% 41.5% 12.0% Fake 8.6% Disqualified

5. Examining Error Modes by Category and Model

Unassisted AI Error Rate by Research Category

Earnings Call Transcripts (Q&A)
100.0%
Qualitative Risks & Litigation
20.0%
Supply Chain & Commitments
18.0%
Sovereign & Export Controls
15.0%
Revenue Segments & Recasts
14.0%
Geographic Revenue & FX Drag
13.0%

Case Study 1: Qualitative Regulatory Inversion (Tesla Supreme Court Ruling)

FORENSIC CASE STUDY

Tesla Supreme Court Tariff Ruling (WS1000-0016)

Prompt Task: In Tesla's Q2 2026 Form 10-Q (Note 11), what was disclosed regarding the February 2026 US Supreme Court tariff ruling?
EXPECTED GROUND TRUTH
Supreme Court issued a ruling invalidating certain tariffs previously imposed under IEEPA, preserving Tesla's tariff refund claims.
Unassisted Parametric Model
Inverted the legal outcome: claimed the Supreme Court upheld tariffs against Tesla.
Massari MCP Layer
Retrieved exact verbatim disclosure from Note 11 with sentence coordinate highlighting in 1.1s.

Case Study 2: Forward Margin Guidance in Earnings Q&A (NVIDIA Blackwell Ramp)

FORENSIC CASE STUDY

NVIDIA Blackwell Gross Margin Guidance in Q&A (WS1000-0009)

Prompt Task: When sell-side analyst Stacy Rasgon (Bernstein) asked CFO Colette Kress to clarify "low-70s" gross margins during the Blackwell ramp, what exact range was given?
EXPECTED GROUND TRUTH
CFO Colette Kress defined "low-70s" as 71.0% to 72.5% before re-accelerating to mid-70s.
Unassisted Codex / ChatGPT 5.5
"Not disclosed by filer in SEC Form 10-Q/10-K. Earnings transcripts are not filed on EDGAR."
Massari MCP Layer
Retrieved speaker-diarized transcript excerpt with exact analyst exchange in 1.2s.

6. High-Entropy SEC Accession Numbers: The Forensic Lie Detector

In fundamental finance, a model that generates fluent, plausible numbers from ungrounded memory creates catastrophic compliance liability under SEC Rule 10b-5.

Every filing submitted to SEC EDGAR is assigned a unique 20-character identifier: 0001045810-26-000052 (10-digit CIK, 2-digit Year, 6-digit Washington Sequence). Because the sequence is issued sequentially by the SEC server at the millisecond of submission, the entropy exceeds $10^{18}$ states.

ITS=(
Verified Grounded PassesTotal Tasks (1,000)
)×(1
Fabricated AccessionsTotal Attempted
)2×100
Where fabricated citations incur an exponential quadratic penalty, reducing model trust score toward zero.

Institutional Trust Score (ITS) with Quadratic Falsification Penalty

Massari MCP Layer
100.0%
ChatGPT 5.5 (Codex Web)
79.9%
Claude Opus (Unassisted UI)
6.4%
Parametric LLM (Raw Memory)
8.6%

7. The Modern Infrastructure Layer: Transcripts, MCP, and Excel Integration

As frontier LLMs master tabular data extraction, the competitive advantage for institutional research platforms has shifted to comprehensive infrastructure:

  1. Speaker-Diarized Earnings Call Transcripts: Real-time semantic search over 19,000+ public companies, capturing Q&A exchanges and guidance nuances excluded from EDGAR.
  2. Universal Model Context Protocol (MCP): Connect your existing AI (Claude Desktop, ChatGPT, Gemini) directly to 20+ years of primary SEC archives.
  3. Live Dynamic Spreadsheet Bindings: Native Excel formulas (=MASSARI.FIN(ticker, metric, period)) that calculate instantly with zero transcription friction.
Massari AI Filing Intelligence: Verified claim accounting with 1-click sentence coordinate highlighting.
The Massari Excel add-in pulling live, cited figures directly into models via live recalculating formulas.

Frequently Asked Questions

What did Daloopa's original benchmark show?

Daloopa's 2024 benchmark evaluated LLMs on 500 fundamental retrieval questions and showed that early models frequently suffered from rounding errors, period shifts (e.g. confusing fiscal vs calendar quarters), and table hallucinations in SEC filings.

How does The Wall Street 1,000 expand upon Daloopa's benchmark?

The Wall Street 1,000 expands the benchmark scale to 1,000 questions across 100 public equities and introduces six deep institutional research pillars, including qualitative litigation, supply chain commitments, sovereign export risks, and 200 analyst-led earnings call Q&A tasks.

Why do general AI models fail on earnings call transcripts?

Earnings conference calls and sell-side analyst Q&A sessions are copyrighted audio events. They are not filed on SEC EDGAR. Without a specialized financial terminal API like Massari MCP, models have zero access to earnings call transcripts and cannot retrieve executive guidance.

How does Massari MCP eliminate financial hallucinations?

Massari MCP connects LLMs directly to 20+ years of primary SEC EDGAR XBRL archives and speaker-diarized transcripts. It provides sentence-level line-coordinate audit trail for every number, ensuring 100% deterministic ground truth.


Access the Benchmark & Platform

To deploy Massari's 36 read-only MCP tools in your research practice:

See the plans · All notes

Book a meeting