Skip to content

Benchmark

FilingFacts benchmark: AI models on SEC filing questions

On 150 questions about SEC filings, gpt-5-mini answered 19.3% correctly on its own and 72.0% with an MCP server that reads the filings.

Each question asks for a figure from a company's SEC filings, a change or ratio between two of them, or a figure the filings do not contain. The expected answer is computed from the filings' XBRL data and rechecked by a separate program. "Closed book" is the model alone; "with MCP" gives it one tool, canli-validation-mcp's company history, which reads the filings.

Accuracy on 150 questions (30 per question type), with 95% intervals
ModelClosed bookWith MCPClosed book: numbers given that were wrongRun records
gpt-5-mini19.3% [13.8%, 26.4%]72.0% [64.3%, 78.6%]100.0% of 17closed run · MCP run
gpt-5.4-mini21.3% [15.5%, 28.6%]66.0% [58.1%, 73.1%]96.4% of 84closed run · MCP run
gpt-520.0% [14.4%, 27.1%]not yet run100.0% of 24closed run
gpt-4.1-mini18.0% [12.7%, 24.9%]not yet run98.4% of 64closed run
gpt-5.412.7% [8.3%, 18.9%]not yet run97.3% of 112closed run

What the runs show

Without the filings, models mostly either decline or state a number, and the numbers they state are almost all wrong. The fourth column counts them. A tool that reads the source turns most of those into correct, checkable answers.

Check it yourself

Every run record holds the exact questions, the model's raw answers, every tool call and the scoring. replay-evaluation.mjs rescores a record offline. The questions are in filing-facts-v0.jsonl, and the harness is open source. All figures on this page are in leaderboard.json.

Limits

  • Each row is one run of one model on a fixed stratified sample of FilingFacts v0 questions; a different sample or a rerun can move a figure within its interval.
  • Closed-book scores measure what a model produces without the filings, not what it knows: most correct answers need a figure from one specific filing.
  • Expected answers are computed from SEC XBRL data and rechecked by an independent program. They have not yet been verified by human experts.