On 150 questions about SEC filings, gpt-5-mini answered 19.3% correctly on its own and 72.0% with an MCP server that reads the filings.
Each question asks for a figure from a company's SEC filings, a change or ratio between two of them, or a figure the filings do not contain. The expected answer is computed from the filings' XBRL data and rechecked by a separate program. "Closed book" is the model alone; "with MCP" gives it one tool, canli-validation-mcp's company history, which reads the filings.
| Model | Closed book | With MCP | Closed book: numbers given that were wrong | Run records |
|---|---|---|---|---|
| gpt-5-mini | 19.3% [13.8%, 26.4%] | 72.0% [64.3%, 78.6%] | 100.0% of 17 | closed run · MCP run |
| gpt-5.4-mini | 21.3% [15.5%, 28.6%] | 66.0% [58.1%, 73.1%] | 96.4% of 84 | closed run · MCP run |
| gpt-5 | 20.0% [14.4%, 27.1%] | not yet run | 100.0% of 24 | closed run |
| gpt-4.1-mini | 18.0% [12.7%, 24.9%] | not yet run | 98.4% of 64 | closed run |
| gpt-5.4 | 12.7% [8.3%, 18.9%] | not yet run | 97.3% of 112 | closed run |
What the runs show
Without the filings, models mostly either decline or state a number, and the numbers they state are almost all wrong. The fourth column counts them. A tool that reads the source turns most of those into correct, checkable answers.
Check it yourself
Every run record holds the exact questions, the model's raw answers, every tool call and the scoring.
replay-evaluation.mjs rescores a record offline. The questions are in
filing-facts-v0.jsonl, and the harness is
open source.
All figures on this page are in leaderboard.json.
Limits
- Each row is one run of one model on a fixed stratified sample of FilingFacts v0 questions; a different sample or a rerun can move a figure within its interval.
- Closed-book scores measure what a model produces without the filings, not what it knows: most correct answers need a figure from one specific filing.
- Expected answers are computed from SEC XBRL data and rechecked by an independent program. They have not yet been verified by human experts.
