The bots that quietly bleed money are rarely the ones with a bad strategy. They are the ones where the backtest and the live runner are two different programs that happen to share a name. The backtest says the strategy makes money. The live account says otherwise. And when you go looking for the reason, there is no single bug to point at. There are twenty small differences in how the two paths handle candles, fills, timestamps, and rounding, and each one is individually defensible, and together they turn a profitable curve into a flat one.
I have watched this happen enough times that I stopped treating it as a mystery. The strategy did not stop working. The code that runs it live is not the code you tested. Those are two claims, and only the second one is fixable.
One engine, two consumers
The cleanest way I know to avoid this is boring, and it is the thing most people skip because it feels like over-engineering when they are excited about an idea. Write the strategy once, as a pure function, and let two different runners call it. The simulator feeds it historical candles. The live runner feeds it real ones. Neither of them contains any strategy logic of its own.
Concretely, the strategy takes some state and a bar of market data and returns a decision. Long, short, flat, size, maybe a stop level. It does not fetch data. It does not place orders. It does not know or care whether it is running against a CSV from three years ago or a websocket feed. All of the environment-specific stuff, where candles come from, how orders get routed, how fills come back, lives in the runners, outside the strategy.
The moment strategy logic leaks into a runner, you have created a place for divergence to hide. Someone adds a little smoothing to the live feed because the raw data looked noisy. Someone special-cases a fill in the backtest to make a trade line up. Six weeks later nobody remembers either change, and the two paths have quietly grown apart. If the strategy is a single function with no I/O, there is simply nowhere for that drift to live.
The candle-close bug that flatters every backtest
The most common way a backtest lies to you is timing, and it almost always makes the strategy look better than it is. It works like this. In a backtest, when you iterate over historical bars, the entire bar is sitting in front of you. Open, high, low, close, all known. It is very easy to write a signal that uses the close of a bar to decide something, and then acts on that same bar. In the backtest that is fine, because the data is already there. Live, it is impossible. At the moment a candle is forming you do not know its close. You find out the close only when the candle is done, which means the earliest you can act on it is the open of the next bar.
So the backtest quietly gives your strategy information it could never have had in real time. Every entry is a fraction of a bar early, every exit is a fraction of a bar early, and on a strategy that trades often, that fraction compounds into a return number that is pure fiction. This is lookahead, and it is the reason a lot of strategies that look incredible in simulation go flat the day they go live.
The fix is a rule you enforce in the engine, not a thing you remember to do. A signal computed from a bar's close may only execute on the next bar's open. If you want to be strict about it, the strategy function should never receive the current forming bar at all. It receives only closed bars, and it emits a decision that the runner executes on the next open. Make the simulator obey the same rule, and the easy lie disappears. Anything else the strategy needs, indicators, moving averages, whatever, gets computed from closed bars too, so an indicator value is never available before the bar it belongs to has actually finished.
Replay live days through the backtester
Here is the part that turns all of this from a principle into something you can check. Record what actually happened live. Every bar the live runner saw, every decision the strategy emitted, every order sent, every fill received, with timestamps. Then take that recorded stream of bars, feed it back through the backtester, and diff the trades the two produced.
If your engine is genuinely shared and your timing rules are honest, the backtest replay of a live day should produce nearly the same trades the live runner produced. Same entries, same exits, same sizes, same bars. When they match, you have earned the right to trust the backtest, because you just proved it agrees with reality on days you have reality for. When they do not match, the diff hands you the exact bar where the two paths disagreed, and that is your bug. No arguing about whether the strategy works. You have a timestamped disagreement to go read.
A parity suite worth keeping tends to check a few things on each replayed day:
- Every trade in live has a matching trade in the replay, same direction, same entry and exit bar.
- Fill prices agree within a tolerance you set on purpose, so you are accounting for real slippage rather than pretending it does not exist.
- Position size and any leverage match to the unit, since rounding and minimum-lot rules are a classic silent source of drift.
- The count of signals emitted matches, so a signal that fired live but not in replay, or the reverse, gets caught immediately.
- No trade in the replay uses information from a bar that closed after the live decision was made, which is your lookahead tripwire.
You will not get a perfect diff, and you should not want one. Slippage is real, a fill will land a tick off, a websocket will drop a message and you will backfill a bar slightly differently. The point is that a tolerance is a number you choose deliberately, and anything outside it is a real finding. Divergence stops being a vibe and becomes a value you can plot over time and watch for spikes.
A workflow you can actually run
The habit that keeps a bot honest is small and repeatable. Every strategy change goes through the shared engine, never into a runner. Before anything goes live, run it through the simulator with the closed-bar and next-open rules enforced. Once it is live, log everything with timestamps. On some regular cadence, replay recent live days through the backtester and diff. When the diff exceeds your tolerance, you stop and read the offending bar before you touch the strategy, because the problem is almost never the idea, it is the plumbing.
None of this makes a mediocre strategy good. What it does is stop a good strategy from dying in translation, which is a much more common way to lose money than most people admit. When we built the backtesting and live execution paths for Blockcircle, sharing one strategy engine between simulator and live runner was the decision that saved the most debugging later, precisely because it removed the places where the two could quietly disagree.
If you take one thing from this, make it the replay diff. It is not glamorous and it will feel like busywork the first few times you set it up. Then one day the diff lights up on a bar you never would have suspected, you go read it, and you find the exact line where your live bot was trading a slightly different strategy than the one you tested. That is a much better afternoon than staring at a flat equity curve wondering where the edge went.