00 / Progress
Built in the open, phase by phase.
Canli Capital is a company you can audit. The engine is 12 phases deep, each tested before the next began. This page tracks what is proven, what is not, and what is next.
12 phases shipped 3,961+ automated tests mypy --strict continuous integration
01 / Corrections
The corrections we have published against ourselves.
This is the list of corrections we have published against our own record, newest first, copied verbatim from the signed transparency log.
-
— the known-open defects list said the sizing overlay's scale defect was open. The fix landed on 2026-08-18, the same day that entry was written, with a test pinning the realized leg to the unlevered book; the leg's survival across production restarts followed on 2026-09-06. The entry understated the code for nineteen days and now records both dates; the estimated drawdown cost while it was open is unchanged.
-
— the /progress roadmap said the next breadth would be managed-futures trend, but that needs futures data we have not yet invested in. That sentence was wrong from the day AlphaTrend, the managed-futures trend sleeve on exchange-traded funds, entered the paper book, and it stayed on the page beside a description that counted four live sleeves. The paragraph now says what is true: four sleeves run on paper and the next admission is a gate, not a schedule, with no new candidate admitted since the pre-registered contract took force. The retired sentence is on the retracted-claims blocklist, so it cannot be republished without this note beside it.
-
— the crypto sleeve did not rebalance for five weeks, and the cause was our plumbing, not the market. Its weekly rebalance failed on 2026-08-13, 2026-08-20, 2026-08-27 and 2026-09-03 with one error: the signal cross-section was built from the shared universe store, which carries the equity universe beside the crypto perpetuals, and the carry feature asked the instrument record for an equity's funding interval on a host whose record holds only perpetuals. The hourly cycles kept marking the book and reporting a healthy hold, the health monitor read only the latest cycle and stayed green, and the positions opened on 2026-07-30 sat unchanged while one short more than doubled against the book. Most of the flagship's forward loss over this window is attributable to that defect rather than to the carry strategy; the marks themselves are genuine, the sleeve simply did not trade when it should have, so the record stands and this note is added beside it. The same audit found two more. The drawdown ladder, the brake that halves and then flattens a sleeve on deep drawdowns, was rebuilt from nothing by every hourly process and so never saw a drawdown at all; replaying the recorded equity curve through it shows it would have halved this sleeve's gross on 2026-08-26, and on restore it enters that half-gross state. And the loop's monthly universe refresh was never wired, so the traded cross-section has been the one selected on 2026-06-01. Fixed today: the signal cross-section is scoped to the sleeve's asset class on both the research and the live path, the ladder restores from the equity curve on boot, and the health monitor now fails on a missed weekly rebalance or on a failed cycle on the rebalance grid. The universe refresh is disclosed, not yet changed: re-wiring it changes the traded cross-section, which is a live-configuration decision and is declared before it ships. Per our append-only posture this correction is added to the record, not swapped into it.
-
— the number we use to justify keeping AlphaTrend does not reproduce, and it overstated the case by about 3x. We publish that AlphaTrend is held 'for measured drawdown reduction (removing it makes the book's max DD 69% worse), not for a demonstrated edge'. Re-measured on the current four-sleeve book, removing it takes max drawdown from -3.68% to -4.51%, i.e. 22.7% worse — not 69%. We looked for a configuration that yields 69% and could not find one: on the three-sleeve book we previously ran at 40/40/20 it is 16.6%, at equal thirds 20.5%, and with the beta overlay applied it falls to 4.4-7.0%. The rest of the sleeve's disclosure stands and is if anything harsher than before: leave-one-out on the current book shows AlphaTrend CONTRIBUTES -0.092 Sharpe (removing it would raise the book from 1.396 to 1.488), its standalone Sharpe of 0.248 is the weakest of the four, and its re-derived DSR is 0.000. So it is held for a real but materially smaller diversification benefit than we claimed. WHAT WE DID NOT DO: change its weight. A weight sweep shows a clean monotone trade — roughly 0.13 Sharpe per 1 percentage point of drawdown, with no optimum anywhere — so any weight we picked off that curve would be a risk preference dressed as a measurement, and choosing the in-sample argmax is the selection trap our deflation discipline exists to prevent. Equal weights remain what the evidence supports: no sleeve's deflated Sharpe justifies more capital than any other's, which is exactly why they are equal.
-
— this site published a 300% one-day gain that never happened, and it stood for three days. Our flagship live curve read 100,000.00 on 2026-08-07 and 400,207.73 on 2026-08-08. ALPHAC is market-neutral and runs gross at or below 1.0x; it cannot quadruple in a session, and it did not. The cause was mundane and entirely ours. On 2026-08-07 the book moved to fresh $1M accounts, and the routine that records broker equity wrote history with INSERT OR REPLACE — merging each account's history into whatever was already stored rather than replacing it. Rows written while a profile still pointed at the SUPERSEDED account therefore survived at any timestamp the new account's history did not happen to cover. AlphaTrend's curve ended up holding $1,000,000 marks from the new account interleaved with a $100,681.45 mark from the old $100k one. The reader rebased on the first mark, turning that row into 10,068, and the step to the next day published as a +893% return; at one third weight that is +297% on the book. Two compounding defects: several marks a day meant two points shared a date, so a 'daily return' was computed between two marks of the same afternoon; and nothing rejected a step that is arithmetically impossible for the strategy. Both are fixed. The writer now REPLACES the curve, since a broker's full history is authoritative on its own and merging into it is unsound — which also self-healed the stored data, taking two sleeve curves from 68 and 52 mixed marks down to 4 and 4 clean ones. The reader keeps one mark per day and refuses to splice a superseded account into a current record. Six tests pin the exact marks that produced 400,207, and one asserts a genuine -10% day is NOT mistaken for an account switch, because a guard that trimmed real drawdowns would flatter this record rather than protect it. The corrected figure for the same window is 99,887.94, i.e. -0.11%. We are stating the wrong number here in full rather than replacing it quietly: it was public, it was flattering, and anyone who looked at this site in those three days saw it.
-
— we have said AlphaTrend 'CLEARED multiple-testing deflation' and it did not. Our deploy gate is DSR >= 0.95. AlphaTrend's DSR is 0.83. **0.83 does not clear 0.95**, and the sentence on this record calling it 'the FIRST sleeve to CLEAR multiple-testing deflation, statistically real, not a backtest fluke' was wrong when we wrote it and has been wrong every day since. What is true is narrower and we should have written it: AlphaTrend scores far better on that measure than anything else we have built, and it was the closest any sleeve had come. That is not the same as clearing. There is a second problem behind the first. AlphaTrend's 0.83 was computed at n_trials=5, while AlphaMax was graded against N=101 on the same record. Those two numbers are not comparable, and presenting them side by side flattered the one graded against fewer trials. Re-graded against the ledger as it actually stands the figure falls, and we are re-deriving it before we publish a replacement rather than quoting a number we have not recomputed. Until that lands, treat AlphaTrend's DSR on this site as WITHDRAWN rather than merely caveated.
-
— for six weeks our two public sites published two different track records for the same book, and the stale one was the flattering one. app.canlicapital.com was serving glass-box artifacts last generated 2026-06-24: a track record reading 0.00% return over '3 days live' with NAV at a clean $100,000, while canlicapital.com served the truth for the identical book, -2.54% over 38 days with NAV $97,459. Anyone who opened the dashboard saw a flat, unblemished record. Anyone who opened the landing page saw a loss. Both were live at the same time. The cause was mundane and that is the point: our publisher copies paper-state.json to both sites but the three glass-box exporters wrote only to the landing site's directory, so the dashboard's copies simply stopped updating on 2026-06-24 and nothing alerted. Our health check verifies that both hosts serve the same paper-state.json; it never checked the glass-box files, so the gap was invisible to the monitor built to catch exactly this. Fixed at the source: all three exporters now write one stamped string to both directories, so the two hosts are byte-identical by construction and cannot disagree even about a content hash. We are stating this plainly because a firm whose entire claim is that its record cannot be quietly re-picked was, for six weeks, publishing two different inception stories on two different hosts, and the more favourable one was the one that was wrong.
-
— the decorrelation we call 'the edge' was measured on a basket containing a strategy we do not trade, and correcting it removes our headline claim. We publish an average pairwise sleeve correlation of -0.04 and describe this book as 'three near-uncorrelated sleeves; that decorrelation is the edge'. That -0.04 was computed over FOUR equity curves, and one of them, prereg_investment, is not a sleeve. The book is and always has been three sleeves: AlphaForge, AlphaMax and AlphaTrend. The script's own comment described its inputs as 'the four curves that make up the live book', which was simply not true of the book. Recomputed on the three sleeves we actually trade, over the identical 728-day window: average pairwise correlation is **+0.072, positive**, not -0.040; average sleeve Sharpe is 0.529; and the worst pair, AlphaMax against AlphaTrend, correlates at **+0.21** because momentum and trend are close cousins and always were. Why this matters more than a third decimal place: a book of N sleeves has effective breadth N/(1+(N-1)*rho), which converges to 1/rho, so its Sharpe can never exceed s/sqrt(rho). At a NEGATIVE rho that ceiling does not exist, and we have been reasoning and planning as though it did not. At the true +0.072 the ceiling is **1.97** — meaning no number of additional sleeves of this quality, not fifty and not two hundred, can take this book past a Sharpe of about 2. Our stated goal is 2.5. On the correlation we have now measured rather than assumed, that goal is unreachable by adding sleeves, and it was our own measurement error that hid the ceiling. Fixed at the source, and the corrected figure is now what the analysis emits. We are stating this in full because it is the single most load-bearing number in our research, we got it wrong, and the error ran in the direction that flattered the plan.
-
— our social card has been advertising a forward Sharpe we stopped believing weeks ago. The image that unfurls whenever this site is shared on LinkedIn, X, Slack or anywhere else read 'FORWARD EXPECTATION 0.7 TO 1.0 SHARPE, AFTER DEFLATION'. Our published expectation is 0.3 to 0.9, so the FLOOR was overstated by a factor of 2.3, and the words 'after deflation' made it a specific technical claim rather than loose marketing. The same card said 'three quant algorithms' when we publish four, and called the book market-neutral without mentioning the +20% net-long overlay we disclose everywhere else. It went stale when we lowered the band on 2026-07-29 and nothing pointed at it, because it is an image rather than a number in an artifact, and our freshness checks only ever look at artifacts. Replaced with a card that carries NO performance figure at all: a picture that travels across the internet for months and cannot be corrected in a reader's cache has no business quoting a Sharpe ratio. Note for anyone verifying: the social platforms cache these aggressively, so the old card may keep appearing on previously-shared links for some time. That is the clearest possible argument for never putting a number on one. Also corrected in the same pass: our app manifest still named the firm 'Canli Capital / AlphaForge' and described it as an engine for crypto perpetual futures, which describes ONE 40% sleeve rather than the book.
-
— we have been counting our own experiments in one place while running them in four, so every deflated Sharpe we publish is more flattering than it should be. Our honesty machinery penalises a result for how many ideas we tried before we picked it, and that count is read from a trial ledger. It reads ONE ledger, var/experiments.jsonl, which holds 127 rows and 101 distinct hypotheses. But research run under other data profiles writes to its own ledger, and those were never totalled anywhere: 35 more rows under the managed-futures profile, 12 under Sharadar, 1 under futures. Deduplicated across all four, the honest count is **127 distinct hypotheses, not 101** — our real search was 26% larger than the number we deflate against. Which directory a trial landed in is a filing convention. Multiple-testing correction does not care about filing conventions: those were all real hypotheses about the same markets, spent by the same researcher, hunting sleeves for the same book. The consequence is one-directional and it is against us: an undercounted N makes a deflated Sharpe look better, so every DSR on this record computed against N=101 (and the earlier N=93, and the N=27 our equity sleeve was first published at) reads better than the truth. We are not restating those figures today because re-deriving each one properly is its own piece of work and we would rather publish the error immediately than a hasty replacement. Treat every DSR on this site as an upper bound until they are re-derived. Fixed at the source: the audit now counts the union of every ledger, which also closes the evasion this revealed — a future search can no longer duck the budget by writing to a directory nobody adds up.
-
— a factual claim in our 2026-07-20 incident note above is now false and we are marking it rather than editing it. That note said 'Bybit responds' as evidence the Binance failure was venue-specific. Re-measured from the same network on 2026-08-06: Bybit does NOT respond, and neither do Kraken or Deribit. Binance, Bybit and Kraken all fail on connection reset in under a tenth of a second and Deribit times out. The conclusion the original note drew is unchanged and if anything strengthened, since the block is clearly broader than one venue, but the supporting fact it cited is no longer true and a reader checking our work today would find it wrong. One venue does answer: OKX responds normally and serves funding rates, 436 perpetual instruments and hourly candles. We are evaluating it, and we note in advance that moving venue is not a configuration change: the same carry signal computed on OKX has a cross-sectional rank correlation of only about 0.64 with the Binance one, so it would be a materially different strategy wearing the same name, and we will not splice the two curves together and call it one continuous record.
-
— our crypto carry sleeve is a strategy that earns funding, and its LIVE paper account has never once booked funding. We found this ourselves and are publishing it before the fix rather than after. The mechanism, stated plainly so anyone can check it: funding is applied in exactly one place in this codebase, Ledger.apply_funding, and it is called from exactly one call site, the BACKTEST engine. The live paper broker's cash moves on one line and only one line, a fill: cash - qty*price - fee. No funding cashflow can reach the live account by any path. The size of what is missing is not marginal. Our committed walk-forward artifact for this sleeve reports funding_net of $19,500 against a total return of $38,236 on a $100k book over 2022-02 to 2026-06: about HALF of everything this strategy has ever earned is the funding it is designed to harvest. Removing it takes the sleeve's backtested Sharpe from 0.653 to roughly 0.30. So AlphaForge's live record has been running against a hard ceiling of about half its validated Sharpe, and every live number we have published for this sleeve should be read in that light. What we are NOT doing: we are not restating the historical live curve. The marks we published were the marks we observed, and this record is append-only, so the fix applies FORWARD from the day it ships and the understated period stays visible in the chain. One number we are deliberately NOT putting on the record yet: our internal estimate of what funding would have added over the live window is dominated by a single name that redenominated roughly 100x inside the window, which makes the naive figure meaningless. We would rather publish no number than a number we cannot stand behind, so that one waits for the corrected accounting. This is the fourth time a defect in our plumbing, not our strategy, has moved a published number, and the lesson we wrote down in July holds: a fix is not done until a test pins the path that actually runs.
-
— a reverse split was mismarked and it cost us real money and a week of half-size trading; here is everything. During the equity sleeve's July drawdown (itself genuine: a momentum junk-rally squeeze, every construction of the factor lost on the same days, and the signal verified alive — longs 97th momentum percentile, shorts 6th), an audit found the nightly simulation had marked a 1-for-20 reverse split (ALIT) at raw prices: it fabricated a -4.95% simulation day that never happened, leaked about -1.45% of real loss into the live account via an oversized short, and falsely tripped the -10% drawdown brake — so the live book traded at roughly HALF its intended size for over a week on a phantom loss. The corrected curve's true worst day is -2.49% and its maximum drawdown -7.1%; the brake should never have fired and releases on the corrected numbers at the next session. Two more defects found and fixed in the same audit: crypto perpetuals had leaked into the equity book's record through Jun 11 (~+$323 of crypto PnL inside the published equity record — the same unscoped-universe class as the 2026-07-05 crypto bug, now guarded at two layers), and the live recipe had drifted from the exact configuration our published evidence blesses (disclosed here; the strategy itself is unchanged). Fixes: split-aware marking with a sanity guard against bogus records, universe scoped to equities, an asset-class guard in the order path, and regression tests pinning the RUNNING path — because this is the third time a data defect dressed up as performance, and the lesson is permanent: it is never the strategy, always the plumbing, and a fix is not done until a test pins the path that runs.
-
— the crypto sleeve was signal-dead, not 'carry compressed'. From go-live (06-23) to 07-05 AlphaForge held cash and we published that funding carry was compressed. An internal audit found the true cause: a wiring bug blended equity-fundamental alphas (undefined on crypto) into the live signal, invalidating it every cycle — the validated carry-only configuration was never what ran live. Fixed same day (the loop now runs the exact blessed walk-forward config: carry_fund_21, weekly rebalance, with per-cycle signal-health logging so a dead signal can never masquerade as a quiet hold again). The $100k cash equity curve was genuine; the published EXPLANATION was wrong. Per our append-only posture this correction is added to the record, not swapped into it.
02 / Edge status
What the build has and has not proven.
Twelve phases produced a system that is hard to fool: leak-proof data, real costs, and tests most strategies never survive. The engineering is proven. A standalone crypto edge is not yet proven; the live record is too young to claim one. A test that comes back empty is still a test that ran.
// A platform that never reports a strategy failing its own bar is not running the bar.
- Daily bars, equity lake
- 24.7M+
- US stocks, survivorship-free
- 6,8351997 to 2026
- Point-in-time fundamentals
- 380K+
- Crypto instrument universe (live + delisted)
- 775
// Running 24/7 on paper across the live and delisted book. No funded performance in play. The build in the open.
03 / The build
The integrity came first.
The order of the work is the thesis. The data integrity, the cost authority, the validation gauntlet, and the crash-safe execution loop were built before any serious effort went into chasing return, because a good backtest only means something on top of a process honest enough to trust. The engine is statically typed under mypy strict, every phase is covered by tests that must pass in continuous integration before the next phase starts, and research and live code share one path. Nothing here is a prototype dressed up as a product.
12 phases / 3,961+ tests, green / mypy --strict / CI on every change
04 / The written record
A build log without post-mortems is a list of wins.
The phases above say what was built. The engineering notes say what went wrong while building it: a position that lost 99 percent and the guard I decided not to add, seventy files published by a check that could not fail, and the arithmetic that decides whether a Sharpe ratio means anything. Each one is bound to a published artifact or a script you can run.
05 / Roadmap
The way forward is breadth.
A single asset class is one bet on one market staying inefficient one way. The answer to a thin standalone edge is more lowly correlated sources of return through the same machine. Four sleeves now run in the book on paper: equity momentum, crypto funding carry, managed-futures trend on an exchange-traded basket, and a macro surprise sleeve. The next one is a gate, not a schedule: every candidate must clear the same pre-registered admission contract, and since that contract took force no new candidate has, so the fifth sleeve stays a research question until the evidence says otherwise. The engine was built multi-asset from the first line: one data contract, one cost authority, one validation library, one execution loop. A new strategy inherits that discipline and clears the same gauntlet before it earns a dollar.