The question arrives in the same form every time. The research said the strategy did one thing, the live account did less, and the allocator wants to know why. What usually comes back is a paragraph containing the words slippage, market conditions and adjustment period, and that paragraph is why the follow up meeting exists. It is not that the explanation is wrong. It is that it cannot be checked, so it reads as an excuse regardless of how true it is.
Replace it with a table that adds up. Backtest gross, four attributed deductions, live net, and a residual line that is explicitly labelled as unexplained. The discipline is arithmetic rather than rhetoric: if the lines do not reconcile to the actual number in the account, you have not finished, and the size of the residual is the honest measure of how well you understand your own execution.
Make the gap reconcile before you interpret any of it
Start from the identity and refuse to editorialise until it closes. Backtest gross return, minus realised fees, minus implementation shortfall, minus the contribution of trades that were blocked or never filled, minus the effect of sizing differences between research and live, equals live net plus a residual. Everything you eventually want to say about signal decay lives in that residual, and you are not entitled to talk about it until the other four lines are computed from records rather than estimated.
Report all of it in basis points on the same capital base, per period, and put the trade count in the header of every column. That last detail does more work than any of the analysis, because it tells the reader immediately how much confidence any line deserves.
One structural rule about the residual: never distribute it. The temptation, once the four lines are done, is to fold the leftover into whichever explanation sounds most plausible. A residual sitting on its own line, labelled unexplained, is the single most credibility building item in the whole table. Allocators have seen a great many decompositions that reconcile perfectly, and they know what that usually means.

Fees are the easy line and still get done wrong
Take fees from statements, never from the schedule you assumed. The two differ more often than people expect, and the differences are systematic rather than random.
The usual gaps are financing and borrow costs that the backtest never modelled at all, per trade minimums that are invisible at research size and material at launch size, venue rebates or tiers that apply to some fills and not others, and accrual timing that puts a cost in a different period from the trade that caused it. On a profile that trades frequently at modest size, the minimums alone can account for a large share of a gap that everyone in the room was attributing to slippage.
Compute the line as realised fees minus the fees the backtest assumed, not as realised fees. The decomposition is about the difference between two worlds, and half the fee number was already in the research.
Shortfall is measured per trade against a stated benchmark
Implementation shortfall only means something once you have said what the backtest assumed it was getting. Name it explicitly in the table's footnote: the close of the signal bar, the open of the next, the mid at signal time. Then measure each live fill against that same reference and express the difference in basis points.
Two presentation choices matter here. Report by instrument, because a portfolio average hides the case where one thin instrument contributes most of the cost while everything else fills fine, and that case has a fix while the average does not. And report a distribution, or at minimum a median and a worst decile, rather than a mean. Shortfall distributions have long tails, the tail is where the money is, and a mean invites the reader to imagine a symmetric error.
This is also the line where you can demonstrate competence rather than assert it. A desk that can produce shortfall by instrument, with a distribution, over a stated benchmark, is visibly measuring its execution. A desk that reports one average number is visibly not.
The trades that never happened
This is the line most decompositions omit, and on a guarded profile it is frequently the largest of the four. The backtest took every signal. Live, some of those signals never became positions, because an exposure cap was already full, a concurrent position limit bound, a filter rejected the instrument that day, or the order simply did not fill.
To value them you need the counterfactual, and the counterfactual has to be built with the backtest's own rules and no hindsight. For each signal the research took and live did not, price the trade at the backtest's assumed entry and run the backtest's own exit logic. Sum the result. That total is the cost, or occasionally the benefit, of your guardrails, and it is a genuinely useful number quite apart from the allocator conversation, because it is the only evidence you will ever have about what your risk limits are charging you.
The practical obstacle is the record. What this line requires is the signals the engine received, not just the orders it sent, and those are different populations. Autopilot describes per source rules and a full audit trail on the profile, and the honest instruction is to open the export and check whether the signals that produced no order are in it. I cannot confirm that they are, and if it turns out that only executed orders are retained, then the parallel record is your job from day one: capture the signal feed independently and reconcile it against the order log each period. That is unglamorous work, and it is the difference between a decomposition and a guess, so decide it before you need it rather than in the week the question arrives.
While you are there, separate the blocks by cause. A cap that bound is a design choice you made and can defend. An order that failed to fill is an execution problem. A signal the engine never received is an infrastructure failure, and it belongs in a different conversation entirely, one about operational risk rather than performance.
What is left is decay, and the sample usually cannot support the word
Once fees, shortfall, blocks and sizing are attributed, the residual is what remains, and it is where the uncomfortable possibility lives. Say the possibility out loud before the allocator does, because they are going to.
The first hypothesis for an unexplained gap is not that the market changed. It is that the backtest was optimistic, and there is a short list of ways that happens: parameters chosen with knowledge of the whole sample, a universe defined using instruments that exist today, an exit that used information available slightly before it should have been, or a research period that simply suited the strategy. Working through that list in front of an allocator is far stronger than being walked through it by them.
Genuine signal decay is the second hypothesis, and it is a statistical claim that most live records cannot support. Before using the word, ask whether the residual is distinguishable from noise given the number of live trades. With forty trades, the answer is almost always no. The correct sentence is that the residual is X basis points over N trades and the sample cannot yet separate decay from variance, along with the trade count at which it will. Naming that threshold in advance turns the next review from an argument into a check.
There is one more line worth adding underneath the table, unprompted, because it will be asked eventually. Which of these gaps is being actively managed and how. Shortfall gets an execution change and a re measurement date. Blocks get a documented decision to keep or widen the cap, with the counterfactual cost as the evidence. Fees get renegotiated or the strategy gets fewer, larger trades. And the residual gets a sample threshold rather than a promise. What that turns the meeting into is a review of four measured processes, which is a conversation you can have every quarter, instead of a defence of a single number, which is a conversation you lose slowly.