Skip to content
Canli Capital

Research

Methodology and questions

82 research documents, 13 subject hubs and 28 published measurements answer these questions in detail. This page answers them briefly, and every answer links to the document where the answer had to be paid for rather than asserted.

What is a deflated Sharpe ratio, and why do you keep failing it?

A Sharpe ratio measures return per unit of risk. It says nothing about how many strategies you tried before you found that one. If you test a hundred variants and report the best, the best is partly skill and largely the maximum of a hundred draws — and the more you tried, the higher that maximum is expected to be even if every single variant is worthless.

The deflated Sharpe ratio corrects for exactly that. It asks: given the number of trials behind this result and how much they varied, what is the probability the true Sharpe is above zero? A figure that looks impressive at one trial is unremarkable at a hundred and fifty. This book's gate is a deflated Sharpe of 0.95, and it is applied against a selection count that is published rather than estimated.

We keep failing it because it is the correct answer. When every historical result was recomputed against the current selection count, 33 variants were restated and 0 of them clear 0.95. That restatement is published in full, variant by variant, with the historical figure beside the corrected one.

The full restatement, every variant The trial ledger the deflation counts against

Why do you publish strategies that failed?

Because the failures are the denominator. A deflated result is only honest if the number of attempts behind it is real, and the only way to make that number checkable is to publish the attempts. 46 candidates have been killed here against 3 that survived, and each kill names the specific number it died on.

There is a second reason, and it is the one that matters to another researcher. A kill log is a table: it tells you the outcome. A paper tells you the reasoning — the universe, the window, the construction, and what would have had to be true for the candidate to live. The outcome saves nobody any work. The reasoning either saves them the experiment or gives them the grounds to show we were wrong.

All 46 killed candidates An example: a pre-registered momentum variant, killed

What is point-in-time data, and why does it change a result?

Point-in-time data is data as it stood on a past date, rather than as it stands today. The distinction sounds pedantic and it is usually the difference between a real result and an imaginary one. Index membership is revised, macroeconomic series are revised, filings are amended, and today's version of the past silently excludes everything that went bankrupt, got delisted, or was corrected.

Two failures follow from getting it wrong, and both flatter. A backtest on today's index constituents is a backtest on companies that survived, which will show a quality effect whether or not one exists. A signal stamped with the date a fact was measured, rather than the date it became knowable, trades on information from the future — and it looks wonderful.

So the vintage discipline is part of the mechanism here rather than a caveat. Where a schedule had to be reconstructed, the provenance of the schedule was audited before the effect was tested, and that audit turned out to be the larger piece of work.

Auditing a schedule's provenance before testing the effect A missing release, found and corrected in public

What does pre-registration actually stop?

It stops the specification moving after the result is visible. A pre-registration fixes the universe, the signal, the horizon, the cost assumption and the pass criteria in writing, before any return data is opened. Once you have seen how a candidate performed, every subsequent choice about it — a different window, a filter, a slightly different construction — is informed by the answer, and the result stops being evidence about the market and becomes evidence about the researcher.

The test of whether a pre-registration is real is whether anything ever dies under it. The registrations here are published alongside the candidates that were registered and then killed, which is the only demonstration that carries any weight.

A locked specification, published before measurement A pre-registered candidate, killed under its own criteria

Why decide whether something can be tested before testing it?

Because opening return data is the expensive step and it cannot be undone. It consumes a hypothesis from a finite budget — 162 of 320 consumed so far — and once a researcher has seen how a candidate performed, that knowledge contaminates every later decision about it.

A feasibility protocol asks three questions first: does the evidence exist in a point-in-time form, can it be extracted at the rate the identity assumes, and are the resulting positions executable at a cost the mechanism can pay. If any answer is no, the idea is not weak — it is unmeasurable here, which is a different and more useful verdict.

The move this guards against is specific. A gate that fails by a few points invites exactly one response: widen the detector until the number clears. That is tuning a measurement to agree with a target. So the ceiling a perfect detector would reach is computed first, and the question is settled on arithmetic rather than on effort.

Asking whether a gate is reachable at all The same question asked of every untouched family What the failures had in common

How can a strategy be real and still not worth trading?

Because the edge and the cost are measured in the same units and the cost usually wins. A signal that predicts a small move over a long horizon can be entirely correct and entirely consumed by spread, fees, borrow, financing and market impact — and none of those show up in a backtest that prices fills at the mid.

The more instructive failure is not modelling a cost at the wrong level but as the wrong kind of thing. Latency represented as a flat basis-point addition is treated as a microstructure effect; an order submitted after the close and filled at the next open is not crossing a spread at all, it is holding unhedged overnight exposure with a fat tail. Getting the size of such a term right does not fix having the wrong term.

Which is why the cost check here published what it could not answer as well as what it could, and concluded that no cost parameter should move on that evidence — the recording schema should.

The modelled cost against what the live book paid How execution is modelled, and its limits

Why is correlation between strategies the binding constraint?

Because a book's Sharpe ratio depends on how many sleeves it has, how good each one is, and how correlated they are — and past a handful of sleeves the correlation term dominates. Adding a tenth strategy that moves with the nine you already own adds turnover and almost no diversification. The arithmetic is unforgiving and it is published rather than described.

The uncomfortable finding is that structural intuition gets this backwards. On the only correlation structure this book has actually measured, the pair sharing a factor family is the most correlated in the book, and the pair sharing an asset class is negatively correlated. Breadth counted in asset classes is not breadth.

The measured structure, and the ordering that follows The book measured with and without one sleeve

Is this real money, and why is there no return to look at?

No. Every sleeve here trades on paper. No real capital is deployed, nothing on this site is investment advice, an offer or a solicitation, and a verified record of a simulation is still a record of a simulation.

The live paper record began 2026-08-07 and has accrued 15 days. That is far too short to say anything about skill, which is stated everywhere the record appears rather than only in a disclaimer. The honest position is that the forward record is the only instrument that defeats the multiple-testing problem, and it defeats it by running for years, not by being presented confidently.

The live record, as it accrues How long it must run before the gap is measurable

Have you ever been wrong, and how would I find out?

Repeatedly, and in the same place and format as everything else. A withdrawn figure, a defect found in one of our own gates, an earlier published answer superseded by a better measurement of the same thing: all of it sits in the same corpus as the results, because a record that contains only its wins is not a record.

The most useful ones are the corrections against our own methodology rather than against a number. A drawdown study was re-run through the estimator production actually uses, and it superseded an answer already on this site. A macro release was found missing and the affected figures were restated in public.

An earlier answer superseded by a better measurement A missing release, restated in public All 28 published measurements

What has to be true before a strategy is added to the book?

A written contract, applied by the production evaluator rather than by judgement. It sets minimum out-of-sample observations, significance floors that survive autocorrelation, a correlation ceiling against the existing book, a capacity floor, and a deflated-Sharpe requirement for the book as a whole. It is published in full, thresholds and all.

The contract has been wrong before, and that is published too: an earlier version contained floors that no candidate could satisfy simultaneously — a gate nobody can pass is not a strict gate, it is a broken one — and each was found by testing the gates against each other rather than against a candidate.

The admission contract in force It run against a real candidate

How do I check any of this without trusting you?

Download the published artifacts and recompute their hashes with nothing but Python; verify the Ed25519 signatures and re-derive all 406 links of the append-only chain behind the track record; clone the engine and run its determinism test. The exact commands are on one page, and so is a section on what each check cannot prove.

That last part is the honest half. A matching hash makes a number unedited, not correct. A valid chain says nothing about what was recorded before the chain began. A deterministic engine is not an accurate one.

The commands, and what they cannot show The glass box itself

Who is behind this?

One person: Arhan Canli, who wrote the engine and every one of the 82 documents published here. There are no credentials on the founder page, deliberately — a credential is a claim you would have to take on trust, and the entire argument of this record is that you should not have to.

The founder page, and the one claim on it you can verify

Where to go next

The research library · Every measurement · How the engine works, stage by stage · How to check all of it

Nothing on this site is investment advice, an offer, or a solicitation. The book trades on paper and no real capital is deployed.