3 / 5 · 7 min

Four ways to fool yourself while testing

A test more often confirms an idea not because it works but because the tester quietly helped it. All four ways look like careful work.

Why this is not about honesty

None of the four requires bad intent — on the contrary, each looks like diligence. You refine parameters so the rule describes the market better. You take current coins because old ones interest nobody. You recall the trades you remember. All of this is reasonable, and that is exactly why the catch is invisible from inside. The only defence is to know these four traps by name and check yourself against each separately.

First: fitting to history

You took a rule, ran it, got 52%. Tried a slightly wider stop — 55%. Used a different averaging period — 58%. A couple more refinements and there is a system with 61% wins. The problem is that a PURELY RANDOM rule produces exactly the same result if you search variants the same way. Measured on synthetic data with a true 50%: the best of ten variants shows 57.7%, the best of fifty 61.2%, the best of two hundred 63.6%. You did not find a regularity, you found the luckiest combination.

What searching gives on random data

Variants triedThe best showsTruth
150.0%50%
1057.7%50%
5061.2%50%
20063.6%50%

The defence has two rules. First: count the variants. If you searched twenty combinations, the winner's result must be read not as 61% but as the best of twenty, which is an entirely different quantity. Second: the fewer tunable numbers a rule has, the less you are able to fit it. A rule with one parameter is hard to fake; a rule with five can be fitted to anything, noise included.

Common mistake

Second: looking into the future

The least visible mistake, because it is technical. You test the rule buy when the daily candle closed above the level — and use the day's closing price to decide whether to enter in the morning. But in the morning you did not know that price. The same family: using the day's final volume, the day's high, a coin that later grew. The sign of the trap is simple: if the decision needs data that appeared AFTER the moment of entry, the whole test is invalid — and its result is usually brilliant, because you are effectively peeking.

Third: survivorship

Testing an idea on today's list of coins means working with those that survived. The dead did not make the list, and there were plenty: our history holds 1423 instruments while the live table now shows 1218 — that is 205, one in seven, that traded and vanished. A strategy of buying dips on altcoins looks wonderful on survivors: the ones that never bounced never reached today's list. An honest test requires taking the instrument list AS OF THAT DATE rather than today — and that is the only way to see how often the dip turned out to be a one-way road.

Common mistake

Fourth: picking after the fact

Testing from memory always gives a good result, because memory is not built like a database. Winning trades are remembered vividly; losing ones come back with caveats — I broke the rule there, it does not count. The caveat may even be true, but it is applied one-sidedly: nobody nitpicks a profitable trade. Hence the only cure is a journal filled in BEFORE the outcome, not after. Everything you recall about your trading without records is an opinion about yourself, not data about the system.

How to check yourself against all four

Before believing a result, ask four questions. How many variants did I try before this one looked good? Did I use, anywhere, data that did not exist at the moment of the decision? Where did the instrument list come from and what happened to those not on it? Did every trade make it into the count, including the ones I explained away as rule violations? Even one uncomfortable answer means the number must be obtained again rather than argued about.

Bar replay closes two of the four traps at once. The future is genuinely hidden, so peeking is impossible; and the decision is written down before advancing, so picking after the fact is excluded. The remaining two — fitting and survivorship — do not depend on the tool and are cured only by discipline: count the variants and take the instrument list as of the right date.

The chart and bar replay
Exercise

Recompute your best idea with corrections

Take the rule you believe in most and honestly recall how many parameter variants you tried before settling on this one. If more than five, use the table above to estimate how much of the result the search alone explains. Then check whether the rule uses any data from the future relative to entry. Finally, count how many trades under this rule you did NOT record because it was not by the system. After three corrections there is usually noticeably less edge than it seemed — and that is a better outcome than learning the same thing through position size.

Check yourself

A rule gave 61% wins. How good is that?

It depends how many variants were tried before it. If you tested one variant and got 61% on a sufficient sample, that is a notable result. If you searched fifty combinations and picked the best, pure noise gives exactly the same: on random data with a true 50%, the best of fifty shows 61.2%. The number alone means nothing without the answer to out of how many was this chosen. That is why the count of variants is recorded alongside the result, just as sample size is.

Check yourself

Why can't a strategy be tested on today's list of coins?

Because that list is made of survivors. Our history holds 1423 instruments and the live table 1218 — one in seven traded and disappeared, and it is not on today's list. Any idea that assumes recovery after a fall will look better than it is on such a list: the examples where recovery never came have been removed from the sample by life itself. The correct approach is to take the instrument set as of the start date of the test and include everyone, even those who later died — otherwise you are testing survival rather than the idea.