A friend sent me a backtest a while back with a cross-validated Sharpe somewhere north of three, and the first thing I asked about was the splits. Five-fold cross-validation, scikit-learn defaults, shuffle on. Which meant the number was fiction, and the annoying part is that everything else about the work was careful. Good features, sensible regularization, honest assumptions about transaction costs. One line of validation code quietly poisoned all of it.
The core issue is that k-fold cross-validation assumes your samples are independent and identically distributed. Financial samples are neither, and the specific way they fail matters, because it decides what the fix has to look like.
Where the leakage actually comes from
Take a common setup. You have hourly bars, features built from rolling windows over the last few hundred bars, and a label on each bar that is the sign of the return over the next 20 bars. The label is the troublemaker here. A label stamped at 10:00 is not resolved until 20 bars later, so it contains information about everything that happened in between. Each sample occupies an interval of time, not a point.
Now hand those samples to standard k-fold. Shuffled folds sprinkle training and test samples right next to each other in time. A training sample from Tuesday morning carries a label computed through Wednesday, and sitting inside that window are test samples the model will be graded on. The model gets to study outcomes from the exact stretch of time it is supposed to be predicting blind. Even without shuffling, contiguous folds leak at every boundary, because the tail of each training block has labels that resolve inside the neighboring test block.
There is a second, sneakier channel. Features on adjacent bars are mostly the same numbers. A 200-bar moving average barely moves from one bar to the next, so a training sample one bar away from a test sample is close to a photocopy of it, usually with the same label. Your model does not need to learn anything generalizable to score well. It can memorize the local neighborhood, and the validation score will reward it for doing exactly that.
The result is a validation estimate that is biased upward, sometimes wildly. I have watched classifiers go from an AUC that looked like a real edge to a coin flip once the splits were done properly, with nothing about the model changing except the honesty of the measurement.
Purging and embargoing, in plain terms
The fix has two parts, and the formal treatment comes from Marcos López de Prado's work on financial machine learning. His book has the proofs and the fancier combinatorial version. The working recipe is short.
Purging first. For each test fold, remove from the training set every sample whose label interval overlaps the test fold's time range. If a training sample's label was still being resolved when the test period started, it goes. If its window touches the test window at all, it goes. To do this you need two timestamps per sample: the moment the features are observed, call it t0, and the moment the label is finally known, call it t1. For a fixed 20-bar horizon, t1 is just t0 plus 20 bars. For triple-barrier labels, t1 is whenever the first barrier gets touched, which varies per sample, so store it instead of assuming it.
Embargoing handles the other direction. Purging clears out training samples whose labels bleed into the test window, but samples that begin right after the test window are still contaminated. Their labels are clean. Their features are another story, since rolling calculations look backward into the test period, and serial correlation makes these samples near-duplicates of the final test samples. So you drop an additional buffer of training samples immediately after the test fold, and that buffer is the embargo.
The recipe
- For every sample, store t0 (when the features are observed) and t1 (when the label resolves). If you cannot write down t1 for a sample, you do not fully understand your own labels yet, which is worth discovering now rather than later.
- Split the data into k contiguous blocks in time order. No shuffling anywhere.
- For each fold, take one block as the test set. From the remaining samples, purge any whose interval from t0 to t1 overlaps the test block. The standard overlap check works: keep a training sample only if its t1 falls before the test start or its t0 falls after the test end.
- Apply the embargo: additionally drop training samples whose t0 lands within the embargo window right after the test block ends.
- Train on what survives, predict on the test block, record the metric. When you aggregate, look at the spread across folds as much as the mean. A strategy that only works in two folds out of eight is telling you something.
A note on tooling. Scikit-learn's TimeSeriesSplit avoids training on the future, which is good, but it does not purge, so the boundary leakage from overlapping labels is still there. Plain walk-forward has the same hole. If your labels take h bars to resolve, the last h bars of every training window are peeking into the test set unless you cut them.
Setting the embargo from your label horizon
The floor is easy. The embargo should be at least as long as your longest label horizon. Fixed 20-bar labels mean at least 20 bars. Triple-barrier labels with a maximum holding period of ten days mean at least ten days, and use the maximum rather than the average, because the leaky samples are precisely the slow-resolving ones.
Above the floor it becomes judgment. Long feature lookbacks and strongly autocorrelated features argue for more, and volatility features are the classic offender, since realized volatility stays sticky for weeks at a time. The default that gets quoted around is roughly one percent of the dataset length, which I treat as a starting point rather than a law. The asymmetry is what matters here: an embargo that is too long costs you a little training data, while an embargo that is too short costs you the ability to trust your own results. Err long.
Two diagnostics I would actually run. First, compute your metric under naive k-fold and under purged k-fold with an embargo, side by side. The gap between the two is a rough estimate of how much your process was flattering you, and it is often the most sobering number in the whole research project. Second, shuffle your labels randomly and rerun the purged version. Performance should collapse to chance. If it does not, there is still leakage somewhere, usually a feature accidentally computed with future data.
One honest caveat. Purged k-fold fixes leakage between train and test, and it does nothing about selection bias. If you try four hundred strategy configurations and report the best purged CV score, that number is inflated by the search itself no matter how clean the splits are. Keep a count of what you tried, and discount accordingly.
When we built the overfitting checks into the backtester at Blockcircle, this was one of the first pieces we wired in, mostly because it kept catching our own strategies before anyone else's. The pattern was always the same: a model that looked brilliant under naive validation and merely okay, or worse, under purged validation, and the purged number was the one that live performance ended up agreeing with.
If you take one action from this, rerun your most recent cross-validated result with purging and an embargo set to your label horizon, and see how much of the edge survives. It takes an afternoon, and whatever is left after that test is the part actually worth your risk budget.