You add a condition to a strategy, rerun it, and the curve gets smoother. The drawdown shrinks. Every instinct says keep it. The problem is that a smoother curve is what you get from almost any filter, including filters that are pure noise, because deleting trades reduces variance by construction. Fewer trades means fewer chances to go wrong, and a chart with fewer wiggles is not the same thing as a better strategy.
The evidence you actually want is not on the curve at all. It is in the specific trades the filter removed, and whether they were disproportionately the bad ones.
Count what the filter deleted, then split it
Run the strategy twice, once with the filter and once without, over exactly the same period and symbol list. Then write down four numbers: trade count before, trade count after, and the win and loss counts in each version. Everything useful follows from those.
Take a worked example. The baseline takes 220 trades, 88 winners and 132 losers, and the winners average plus 4.1 percent against losers averaging minus 2.0 percent. Per trade that is 0.4 times 4.1 minus 0.6 times 2.0, which comes to about plus 0.44 percent expectancy.
You add a trend filter and the count drops to 130. So the filter deleted 90 trades. The question is what those 90 were. Say the filtered version now shows 58 winners and 72 losers. That means the deleted 90 broke down as 30 winners and 60 losers, a two to one ratio of losers to winners in the discarded pile, against a roughly three to two ratio in the original population. The filter is doing real work, but less than the curve suggests.
Now redo the expectancy on the survivors. Winners are 58 of 130, so 44.6 percent, and if the average win and loss sizes held roughly steady the expectancy goes to about 0.446 times 4.1 minus 0.554 times 2.0, which is around plus 0.72 percent. Better per trade. But you now take 130 trades instead of 220, so total expected return across the period drops from about 97 units of percent to about 94. The filter improved the quality of each trade and reduced the total, and those two effects nearly cancelled.

That last calculation is the one people skip. Per-trade expectancy going up is not a reason to keep a filter on its own. You have to look at expectancy multiplied by trade count, because a strategy that makes 0.72 percent 130 times is barely different from one that makes 0.44 percent 220 times, and the second one gives you more information and more chances to compound.
The ratio that actually decides it
The clean way to state the test: of the trades the filter removed, what fraction were losers, and how does that compare to the loss rate in the whole population?
In the example above the population was 60 percent losers and the deleted pile was 67 percent losers. That is a lift, but a modest one. A filter that deletes 90 trades of which 80 are losers is a genuinely different animal, and you can see it immediately in this framing while the equity curve would look only slightly better than the mediocre case.
Set yourself a threshold in advance, before you look. Mine is that the deleted pile has to be at least fifteen percentage points worse than the base rate, and it has to hold up when I split the sample in half by date. Anything under that I treat as noise, because a filter that is barely better than random at sorting trades will not stay barely better out of sample. It will land on the other side of the line about as often as not.
The degenerate case, where the filter just deletes the sample
The version to be genuinely afraid of is the filter that cuts 220 trades to 18. The curve looks magnificent. Every one of those 18 trades came from a period where the setup was working, and you have not built a filter, you have built a description of your best trades.
Put a hard floor on it. If a filter takes the strategy under roughly 30 trades in the test window, I stop reading the results, because at that count the confidence interval around the win rate is wide enough to contain both a great strategy and a bad one. A useful gut check: at 30 trades, a run of five straight losses is unremarkable, and a run of five straight losses in live trading is enough to make most people abandon the system. If the sample cannot distinguish those two worlds, it cannot support the decision.
The other tell is concentration in time. Check when the surviving trades happened. If the filter kept 40 trades and 28 of them fall in one six-month window, the filter is selecting for a regime, not a condition. That may still be tradable, but you need to be honest that you are making a bet on that regime returning rather than on a rule that generalises.
Two cheap tests that catch a fake filter
Before you keep any filter, run these. They take minutes and they have saved me from several confident mistakes.
- Invert it. If the filter is real, the trades it excludes should perform noticeably worse than the ones it keeps. Run the strategy on only the excluded trades. If that version is roughly as profitable as the filtered one, your filter is sorting randomly and the smoother curve came from taking fewer trades, not better ones.
- Shift the threshold. If your filter is volume above the 20-period average, try 15 and 25. A real effect degrades gently as you move the number. A fitted one falls off a cliff on one side, which tells you the specific value was chosen by the data rather than by a reason.
The inversion test is the more powerful of the two and almost nobody runs it. It is the difference between "this condition is associated with better trades" and "this condition happened to be true during the good months".
Reading the deleted trades one by one
Numbers aside, spend twenty minutes actually looking at the 30 winners the filter threw away. This is where you find out whether you want the filter regardless of the arithmetic.
Sometimes those winners share an obvious character. They are all the fast ones, the trades that ran 8 percent in three days, and your filter is systematically excluding the tail that carries your best months. A filter that trims small losers at the cost of your largest winners is a bad trade even when the expectancy arithmetic looks flat, because the tail is where the year gets made.
Sometimes the opposite. The deleted winners are all marginal, plus 1 or plus 2 percent, and the deleted losers include three of your five worst trades. That filter is worth keeping even if its effect on average expectancy is small, because it is reshaping the distribution in the direction you want.
The decision rule I use comes down to this. Keep a filter when the deleted pile is clearly worse than the base rate, when the survivors are numerous enough to trust, and when the winners it removed are not the fat ones. Two out of three is not enough. And when a filter fails only on the third test, the useful response is usually not to drop it but to loosen it, so it trims the marginal setups without touching the trades that pay for everything else.