Skip to content

Engineering notes

What went wrong, and what the arithmetic actually says.

Post-mortems on real incidents in this system, derivations of the statistics it gates on, and design arguments about the parts that were hard. Written for readers who want the mechanism, not the summary. Every figure is bound to a published artifact or a script you can run.

  1. 11 Was the winning strategy skill or luck? I ran 399 textbook trading strategies on the same prices and asked whether the winner was skill. The answer depended almost entirely on which test I used, because the tests were answering different questions, and only one of them was the question I meant. · 3 min
  2. 10 The same SEC filing has two acceptance times EDGAR's submissions API labels each filing's acceptance time as UTC. I checked 450 of them against the filing index pages: in 13 of 25 companies' files every time was late by one US Eastern offset, and the same filing read differently depending on which file it came from. · 4 min
  3. 09 My agent dropped numbers when it copied data into a tool call Same model, same question, same tool. With 756 daily returns pasted into the prompt, the agent got 4 of 9 runs right; when the tool read the same series from a file, it got 8 of 9, with about a quarter of the tokens. · 3 min
  4. 08 "Silver price" returned US CPI A tool that turns plain words into a data series picked any series sharing half the words of the request. One generic word was enough, so "silver price" came back as consumer prices, with real numbers and no error. · 3 min
  5. 07 Your backtest beat a t-test. Would it beat a placebo? In 300,000 simulated tests on markets with nothing to predict, a t-test on the best rule of a grid said "edge" between 10.9% and 78.9% of the time at a nominal 5%. A placebo test, which reruns the whole pipeline on the same returns in a random order, stayed between 4.6% and 5.5%. · 4 min
  6. 06 The split that ran backwards From an August repair until 2026-09-14, my engine adjusted almost every stock split in the wrong direction, and every test passed. The data was right by its schema, the code was right by its tests, and the two disagreed about what one number meant. · 4 min
  7. 05 The one symbol that explained the whole gap I could not reproduce my own result. The replay came out materially worse over the same window, on the same code path, from what I believed were the same inputs. The cause was a single delisted token that had quietly left the universe between the two runs. · 5 min
  8. 04 The arithmetic of not fooling yourself A Sharpe ratio of 1.14 with a p-value of 0.017. Publishable, on the face of it. Whether it means anything depends entirely on a number that does not appear in it: how many things you tried before this one. · 6 min
  9. 03 Seventy files that were never meant to be public I published two open-source repositories with a check that proved they were faithful copies. The check passed every time I ran it. It was structurally incapable of failing. · 4 min
  10. 02 The trade that lost 99 percent, and the guard I did not add A carry signal bought a token at 16.15 and closed it at 0.1553. The strategy was not broken. The interesting decision came afterwards, and it was to change nothing. · 5 min
  11. 01 What competitive programming actually bought me, and what it did not There are no segment trees in this trading engine. I want to explain why that is the right answer, and what the contest habit did pay for instead, because the honest version of this story is more useful than the flattering one. · 6 min

The research record, with hash-bound documents and citation metadata, is separate and lives at research. The code these notes describe is at engineering.