Module 1 of 8 · Free to read and share

Why your backtest lied to you

It worked for three years of history. It made money every month you checked. Then you traded it, and it stopped. Nothing was broken, nobody cheated, and the reason is arithmetic you were never shown.

Almost everyone who trades has had this exact experience. You find a rule — maybe you built it, maybe you paid for it, maybe someone on Instagram posted it. You check it against history and it looks superb. You put real money behind it. Within a few months it has quietly stopped working.

The usual explanations are that the market changed, that you lacked discipline, or that you need better entries. Occasionally one of those is true. Usually the answer is simpler, stranger, and far more useful to understand: the strategy was never working in the first place. What you saw in the backtest was real in the sense that the numbers were correct. It just was not evidence of anything.

By the end of this page you will be able to work out, for any strategy you own or are being sold, the exact number it has to beat before it means anything at all. Most strategies do not beat it. Knowing how to check takes about two minutes once you have seen it done, and it will save you years.

I am not going to assume you know any statistics. Where a term is needed, I will explain it in ordinary words first.

A room full of coin flippers

Forget trading for a moment.

Put a thousand people in a room and give each one a coin. Every morning they flip it. Heads, they buy. Tails, they stay in cash. That is the entire strategy. No charts, no news, no skill — we handed them a coin.

Run it for two years. Then ask the room: who did best?

One thousand people. Zero skill between them. One of them now has a two-year track record that looks like genius.

Somebody always wins. Not because they were good — we know none of them were good — but because when a thousand people do something random, the luckiest one looks remarkable. Their equity curve goes up and to the right. Their win rate is excellent. If they posted it online, you would follow them.

And here is the part that should stop you: that person cannot tell either. From the inside, a lucky streak feels exactly like skill. They are not lying to you when they say their method works. They genuinely believe it, and they have the screenshots.

You are not looking at successful traders. You are looking at survivors — and the selection happened long before you found them.

This is not a thought experiment

A thousand is too small. The real retail trading population is in the millions, and the measured outcomes look exactly like the coin room.

Three independent studies, different countries, same answer.
StudyWhoOutcome
Chague, De‑Losso & Giovannetti (2020)Brazilians who kept at it over 300 days97% lost money
Barber, Lee, Liu & OdeanTaiwan, 14 years of records<1% reliably profitable
ESMA broker disclosuresEU retail accounts74–89% lose

Now put those two facts side by side. Almost everybody loses. And the handful you can actually see — the ones with the courses, the signal groups, the followings — are drawn from the luckiest tail of a population of millions.

That is not an accusation of fraud. It is just what selection does. A person with a genuine five-year winning record and no skill whatsoever is the expected output of a million people trying. You would need several such people before the result was even mildly surprising.

Now the uncomfortable part: you did the same thing

You think you tested one strategy. You almost certainly did not.

You tried a 20-day average. Mediocre. So you tried 50. Then you thought it needed a stop, so you added one at 5%, then 10%. Then you decided the asset was wrong and tried a different one. Then you noticed that 2019 was ruining everything, so you started the test in 2020 instead.

Every one of those was a look. And each look is another coin flipper in the room. By the time something finally worked, you may have taken forty looks — and the best of forty random attempts looks good, for the same reason the best of a thousand coin flippers looks like a genius.

This was not cheating. You did what any sensible person does: you kept adjusting until something worked. The adjusting is the problem, because the bar for "this is real" rises every single time you look, and nobody ever told you to raise it.

Three ways people take far more looks than they think

  1. Settings. Trying lookbacks of 10, 20, 50, 100 and 200 is five looks, not one. Five lookbacks × three stop levels × two assets is thirty.
  2. Moving the dates. Starting later because the early years "aren't representative", or stopping before a bad patch, is a look. It is also almost never counted.
  3. Buying someone else's looks. This one catches everybody. If you take a strategy from a course, a paper or a video that tested two hundred ideas and showed you the best one, you inherit all two hundred. Somebody else did the searching on your behalf. Your "one test" is the two hundred and first.

That third one is why buying strategies rarely works, and it has nothing to do with whether the seller is honest. An honest seller who tested two hundred ideas and sold you the winner has sold you the best of two hundred coin flippers.

Three words, in plain English

To check a strategy properly you need three terms. They sound technical and they are not. Read these once and the rest of this page is easy.

The only vocabulary you need

Sharpe ratio
How much you earned, divided by how wildly the account swung around to earn it. Two strategies that both made 20% are not equal if one was smooth and the other was terrifying. Above 1.0 is genuinely good. Below 0.5 means the returns were mostly noise. Your backtesting software already reports this.
Median
The middle value when you line results up smallest to largest — half above, half below. It is more honest than an average, because one spectacular result can drag an average upward while the typical case stays poor.
Noise floor
The score that pure luck is expected to produce, given how many times you looked and over how long. This is the number almost nobody calculates, and it is the whole point of this module. Beating it is the minimum for a result to mean anything.

The table that settles arguments

Two things drive the noise floor, and both are intuitive once said out loud.

More looks, higher bar. Every extra attempt gives luck another go, so the best of many attempts scores higher than the best of a few.

More years, lower bar. Luck cannot keep it up. A fluke that survives six months is easy; one that survives twenty-six years is almost impossible. Time is the only thing that genuinely defeats this problem — which is why long history is worth more than clever ideas.

Find how many looks you took down the side, and how many years you measured across the top. The cell is the Sharpe ratio your strategy must beat.

The noise floor. Read your own row honestly — most home testing happens in the top-left corner, where the bar is brutal.
Looks you took2 years4 years9 years26 years
1 — decided in advance0.000.000.000.00
50.820.580.390.23
20 — a normal afternoon1.320.930.620.37
1001.771.250.840.49
1,000,000 — the whole room3.442.431.620.95

Look at the highlighted row. Twenty looks over two years demands a Sharpe of 1.32. Twenty looks is one focused afternoon of tweaking settings. If your promising strategy scores 0.9, it is not a weak edge that needs refining. It is below what pure luck produces under the same amount of searching.

Now look at the top row. One look has a bar of zero. If you decide the entire rule in advance, write it down, run it once and accept the answer, there is no luckiest-of-many to beat. That is not a loophole — it is the reward for committing before you look, and it is the single most valuable habit in this course.

And compare the bottom row across: a million looks over two years needs 3.44, the same million over twenty-six years needs 0.95. Identical luck, identical searching — the bar falls by nearly four times purely because somebody measured for longer.

Why this is good news, not bad news

This arithmetic is the reason most strategies for sale are worthless — and it is also the only tool that would let you recognise a real one. Right now you have no way to tell a genuine edge from a lucky one, which means every decision you make about strategies is a coin flip of its own. Once you can calculate the floor, you can stop paying for noise, stop grinding on dead ideas, and spend your effort where there is actually something to find.

What it looks like when it happens to you

Theory is easy to nod along with, so here is the same mistake with real numbers. Mine.

There is a well-known paper that publishes 101 trading signals as explicit formulas — no interpretation needed, just code them and run. Ideal test material. I implemented 63 of them.

One stood out immediately. 22.64% a year. A clean formula from a published paper, a return that would embarrass most funds. For about an hour I thought I had found something real.

Then I looked at the other 62.

The median Sharpe across all 63 was −0.13. The typical signal in that set lost money.

So the winner was not a discovery standing in a field of discoveries. It was the best of 63 attempts drawn from a pool whose middle was slightly worse than useless — which is precisely the coin-flipper problem, except the flippers were formulas. Measured against the bar for 63 looks, it did not clear it. It then died a second time on trading fees.

Had I reported only the winner, I would have published a gold mine that did not exist. And I would have believed it, because the number on my screen was genuinely what my code produced.

The rule this produced

Report the distribution, not the winner. If somebody shows you one strategy, ask how many they tested. If they tested sixty and showed you one, you are looking at a maximum, not a result. Ask for the median and the range. The answer to that single question disqualifies most things that get sold.

It is also the question almost nobody asks, which is why it works so well.

The part that is genuinely hard to hear

I tested 479 strategies over about a year: academic papers, factor libraries, Quantpedia entries, Instagram reels, and plenty of my own inventions. One survived everything.

For a long stretch, every single result I produced sat below its own noise floor. Not most of them. All of them. The best thing on my list at one point scored 0.75 against a bar of 1.42.

That is not a story about bad strategies. It is a story about a search that had eaten its own evidence: I had examined the same few years of data hundreds of times, so no result from that data could clear the bar any more, no matter how good the idea.

Two honest conclusions follow, and they point in opposite directions.

Be far more sceptical of your backtest than you are. The floor is usually above your result, and you have taken more looks than you counted.

But a floor describes your evidence, not reality. A strategy with a true Sharpe of 0.3 is serious money at scale, and it is completely invisible at a bar of 1.42. Failing to detect something is not the same as it not being there. The right response is not to give up — it is to stop interrogating the same data and go get evidence the floor cannot swallow: longer history, different markets, and results from time that had not happened yet when you picked the strategy.

That is what the remaining seven modules are. Where the long history is and how to get it free. The one cheap test that does most of the actual killing. Fourteen bugs that made my own results look better than they were, with what each cost. And an honest answer to the question this forces on you: what do you do when a year of work produces one survivor?

Try it on your own strategy

  1. Count your looks honestly. Every setting you tried, every asset you swapped, every start date you moved, every version you abandoned. If the idea came from a course or a video that surveyed candidates, add theirs. Write the number down before you continue.
  2. Count the years you actually measured. Not the data you downloaded — the span you scored.
  3. Find the cell in the table above and compare it with your Sharpe ratio.
  4. Answer one question in writing: if my result is below the floor, what would I need in order to find out whether this strategy is real anyway?

Most people find their result sits below the bar. That is the normal, expected outcome and it is not a reason to stop. It is the reason the rest of this exists — because once you can see the floor, every method for getting evidence over it becomes worth learning.

The full course

One
Survivor

You get the two things that took longest to build: a map of 20+ free data sources with working download scripts, and the scoring tool that calculates every number in this module for you. Plus eight modules explaining how to use them.

DATA
20+ free sources with fetch scripts. Minute bars to 2016, currencies to 2008, company fundamentals, auction calendars, volatility indices to 1990, factor data to 1926 — and the specific trap hidden in each one
TOOL
The scoring kit: noise floors, the placebo test, and a checker that recomputes your results from scratch so a wrong number cannot reach your conclusion
LEDGER
All 479 tests with verdicts and reasons — searchable, so you never pay to rediscover a dead end

The strategy that passed stays mine. The 478 failures are what teach you, and they cost me a year.

01
Why almost everything fails — you just read it
02
The free data map — 20+ sources, decades of history, and the trap in each
03
Where strategies come from, and five ways to kill one in ten minutes
04
Deciding in advance, so your bar stays at zero
05
What to compare against — the mistake nearly everyone makes
06
The one test that killed 9 of my last 10 ideas
07
Fourteen bugs that flattered my results, and what each cost
08
How to judge, and what to do when nothing passes

€400

One payment. Lifetime access. Includes the scoring tools, the data map, and the full ledger of all 479 tests with verdicts.

Get the course

30-day refund, no questions asked. If it does not change how you test, you should not have paid for it.