Skip to content

canli-validation-mcp / Hosted 0.10.1

Check the research behind a backtest

Test selection bias, overfitting and track-record maturity from an AI assistant. Inspect the inputs, calculation limits and reproducible validation receipts.

15 tools in the hosted release. MIT source. Hosted and npm versions are distributed separately.

When to use this server

Use this server when you have strategy returns or summary statistics and need to assess the evidence behind a reported result. Deflated Sharpe accounts for the declared search size; CSCV uses the full matrix of tried variants; minimum track record length asks how much evidence an observed Sharpe needs. The server also checks paper-record disclosures and portfolio breadth.

Connect the hosted server

Paste this URL into a client that supports MCP over Streamable HTTP:

https://canlicapital.com/mcp

For Claude Code, register the same endpoint:

claude mcp add --transport http canli-validation https://canlicapital.com/mcp

The tool list below describes hosted version 0.10.1. The canli-validation-mcp npm package can have a different version; check its published version and README before using a local install. The developer setup guide covers clients, API keys and local validation.

A useful first workflow

  1. Count all variants tried, including failures. Keep the observation frequency and cost assumptions with your returns.
  2. Choose the test that matches your inputs. Send every variant to CSCV or the data-snooping tests; send summary statistics or one return series to deflated Sharpe.
  3. Inspect the result and its limitations. A remote validation returns a receipt that can be retrieved and reproduced from the named source core.

Example tool arguments

Illustrative statistics for a five-year daily record after twenty declared independent trials. Supply your measured statistics and justified search count; this example is not a strategy result.

{
  "name": "validate_deflated_sharpe",
  "arguments": {
    "observed_sharpe_annualized": 1.5,
    "observations": 1260,
    "periods_per_year": 252,
    "skew": -0.3,
    "non_excess_kurtosis": 4,
    "effective_independent_trials": 20,
    "cross_trial_sharpe_sd_annualized": 0.5
  }
}

This is a tool name and arguments for your MCP client, not an HTTP request to paste into the endpoint.

Tools in the hosted release

canli-validation-mcp 0.10.1 tool reference
ToolPurpose and input scope
get_keyIssue a free validation key for this session. Rarely needed: the first validation issues one itself unless CANLI_KEY or local mode is set, and the read tools need none. Quotas: 1000 validations per key per UTC day, 5 keys per client per UTC day, 1048576 bytes per validation request, 1024 bytes per key revocation request, 20000 observations per series, 200 variants per matrix.
validate_deflated_sharpeDeflated Sharpe ratio: the probability (0 to 1) that the selected strategy's Sharpe beats the best that luck gives across the variants tried, with the probabilistic Sharpe and that luck benchmark. Send the seven statistics or a return series. With every variant's returns use validate_overfitting; luck as a trial count, validate_luck_trials; a multiple-testing haircut, validate_haircut_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
validate_overfittingProbability of backtest overfitting (0 to 1) by CSCV: how often the in-sample best variant falls below the out-of-sample median. Needs every variant's returns (periods by variants); with summary statistics only, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
validate_reality_checkData-snooping tests on every variant a search tried: Hansen's SPA p-value that the best beat the benchmark only by luck, White's Reality Check, and the variants Romano-Wolf StepM finds better. Send all variants tried, not only the winners. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
validate_paper_evidenceWhether a paper or simulated performance record meets canli.paper-evidence.v0, with a JSON pointer per failure. Checks structure and required disclosures, not whether the returns are good. This verdict is about the series exactly as submitted. The service never saw the data source, its costs, survivorship, or any lookahead in how the series was built.
validate_breadthBook Sharpe ceiling from adding sleeves of this quality and correlation, the Sharpe at a sleeve count, and the sleeves a target needs. For portfolio construction; it validates no single strategy. This verdict is about the series exactly as submitted. The service never saw the data source, its costs, survivorship, or any lookahead in how the series was built.
validate_track_recordMinimum track record length (observations and years) for an observed Sharpe to beat a benchmark at a confidence level; with observations, the record's probabilistic Sharpe so far. For live or paper records; to size a backtest for its trials, use validate_backtest_length. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
validate_backtest_lengthMinimum backtest length (years) before the best of N independent trials is not expected to reach a target Sharpe by luck; with backtest_years, the most trials those years allow. For planning a search; once it has a result, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
validate_haircut_sharpeHaircut Sharpe for multiple testing (Harvey and Liu 2015): the Sharpe a single test would have needed, by Bonferroni and independent tests, and with the other tests' Sharpes, Holm and BHY. For the probability the Sharpe is real, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
validate_luck_trialsHow many skill-less strategies a search would need for its best to reach this Sharpe by luck (Monte Carlo), and with a trial count, the chance it did. States luck as the best of N random tries; for the probability the Sharpe is real, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
audit_backtestOne-call audit of a strategy's returns: deflated Sharpe, minimum track record and, with every variant's returns, the probability of backtest overfitting, each the matching validator's result with its own receipt. Point returns_file at the backtest's CSV or JSON instead of pasting long series. Prefer it to calling the validators one by one; one validation per check. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
get_receiptFetch a stored verdict by receipt id to re-read it. No key; verify_receipt checks it is genuine. The receipt is content-hashed, reproducible from the open-source core it names, and signed with Ed25519 by a key published at https://canlicapital.com/.well-known/canli-receipt-keys.json.
verify_receiptVerify a receipt offline: its Ed25519 signature against the bundled canlicapital.com key, its output hash and its id. Send an id to fetch it first, or the receipt itself. The receipt is content-hashed, reproducible from the open-source core it names, and signed with Ed25519 by a key published at https://canlicapital.com/.well-known/canli-receipt-keys.json.
service_statusWhether the validation API is up, with its quotas; check after a timeout before resubmitting. No key. This verdict is about the series exactly as submitted. The service never saw the data source, its costs, survivorship, or any lookahead in how the series was built.
company_financial_historySEC-reported financial history for one company in the canlicapital.com reference, by cik or ticker: without a concept, the histories available; with one, observations newest first with accession, form, filed date and unit, plus the source's SHA-256. For point-in-time values use canli-fundamentals-mcp. Public company accounting reference, not market prices, returns, an investment recommendation, or ALPHAC performance. Validate a separately constructed return series with the validation API; accounting values are not returns.

Evidence limits

The tests evaluate exactly the submitted inputs. They cannot establish clean source data, absence of lookahead, complete costs, live profitability or investment suitability. A signed receipt authenticates the service's calculation, not the financial outcome. Local mode is available in the npm package; remote validations store receipts and use the disclosed free-key quotas.

Source for this tool list: released commit 4f7838f8db35 and the machine-readable discovery record. The page follows the same release pin as the hosted handler. Future package work is described in the repository and is not added here until the hosted release changes.

Continue with the underlying evidence

Compare the finance MCP servers or use the developer API and open repositories.