A course in testing trading strategies
I spent a year building a testing process that rejects almost everything. It rejected 478 of 479 ideas — academic papers, Quantpedia entries, Instagram reels, my own inventions. This course is that process, and every one of the 478 failures is a worked example in it.
Take a million people flipping coins to pick trades. The best of them will show a t-statistic around 5.3 — the conventional threshold for a discovery is 2.0. That person has zero skill and a track record that looks like genius.
This is why your backtest looks good. Not because you cheated, but because you looked more than once, and the bar rises every time you look. Almost nobody testing strategies at home knows how to calculate where that bar actually sits — so they find "edges" that are the expected best of however many things they tried.
A third of arbitrary date windows are profitable. So is your edge. The question was never "does this make money" — it is "does this beat an arbitrary version of itself?"
That single test — the placebo sweep — accounted for 9 of my last 10 rejections. It costs one loop of code. Module 6 is entirely about it.
Every bug below is one I actually shipped, caught, and measured. All of them made results look better, never worse — which is why they are so hard to notice. This is the module people tell me they wish they had read first.
| The bug | What it cost |
|---|---|
| Exits resolved from the wrong end of the bar. A dip entry whose target sits above the level price fell to, so bars that rallied then dipped scored as winners. | 0.87 Sharpe PASS → FAIL |
| Weekend returns deleted instead of carried. Aligning a 365-day crypto series onto stock-market sessions throws away every Saturday and Sunday. | 31 points of CAGR |
| Bootstrapped the wrong statistic. Resampled the Sharpe of the difference instead of the difference of Sharpes. The confidence interval did not contain its own point estimate. | PASS → FAIL |
| Sharpe on raw instead of excess returns. A near-zero-volatility cash leg inflated every low-exposure benchmark — one scored 1.59 where the honest figure was 0.12. | 1.47 Sharpe of illusion |
| A data vendor selling the wrong index. A file labelled DAX was the EURO STOXX 50 for 42 straight months. Caught only by a 271% one-minute "move". | whole test invalid |
| A regulator changed a form label. The filing code changed, and a parser keyed to the old string returned zero rows for recent years without raising an error. | silent zero data |
| Drawdown compared with the wrong operator. Drawdowns are negative; the comparison was written the intuitive way, and a 3.8-point-deeper drawdown was reported as passing. | reported a loss as a win |
| Timezone-aware timestamps silently converted. Taking raw values off a tz-aware column shifts everything to UTC and the labels stop matching. Twice. | two broken reindexes |
Six more in the course, including the one where every rebalancing frequency returned an identical number and it took a week to notice that was impossible.
Eight modules, in the order you actually need them. Every module ends with something you can run, not just something you have read.
Best-of-N selection, and how to calculate the noise floor for your own search. The maths of why looking twice raises the bar, and why an inflated floor is as dangerous as no floor at all.
Where real strategies come from — papers, factor libraries, forums, social media — and how to reject most of them in ten minutes without writing a line of code.
Twenty-plus sources I actually pulled, what each one gives you, what it costs, and the specific trap inside each. Minute bars back to 2016, FX back to 2008, fundamentals from the regulator, auction calendars, factor libraries to 1926, options greeks, volatility indices to 1990. Nearly all of it free.
Writing down the rule, the benchmark, the number of variants and the expected result before running anything. The discipline that makes a wrong prediction valuable and a right one credible.
Buy-and-hold is almost always the wrong comparison. How to build an exposure-matched blend — the dumbest version of the same bet — and why one of my candidates beat buy-and-hold in 7 of 7 windows and lost to the blend in 0 of 7.
The highest-value test in this course and the cheapest to run. Rebuild your rule at twenty arbitrary offsets and ask where the real one ranks. It kills whole categories of calendar, session and level claims in one loop.
Fourteen real bugs, each with the measured cost of not catching it, and a test that would have. Look-ahead, silent type coercion, survivorship, the double-subtracted risk-free rate, and the verification that verified nothing because it re-read the same file.
Seven binary conditions, the boundaries that stop you gaming them, and the honest answer to the question this course forces on you: what do you do when a year of work produces one survivor?
If you want someone to tell you what to buy, this is the wrong product and I would rather you knew now. This teaches you to disprove ideas, which is a slower and considerably less exciting skill. It is also the only one that compounds.
Full course — 8 modules, code and data included
€400 one payment, lifetime access, free updates
| In the box | What it is |
|---|---|
| 8 written modules | The method, in the order you need it |
| The criteria document | 7 binary conditions and the boundaries that stop you gaming them — the same one I am bound by |
| The scoring toolkit | Noise floors, placebo sweeps, bucketed information ratio, and a harness that recomputes your results from positions and prices alone |
| The data map | 20+ free sources with working fetch scripts and the specific trap in each |
| 16 pre-registrations | Real ones, written before results, including the predictions I got wrong |
| The full ledger | All 479 approaches, each with its verdict and the reason — searchable |
| The N ledger | How to track what each dataset has already been asked, so your bar rises honestly |
30-day refund, no questions asked. If it does not change how you test, you should not have paid for it. Email and it is refunded.
Instant download — eight modules, the data scripts, the scoring tools and the full ledger, as one file. Card and PayPal. VAT handled at checkout. Updates free for life.