A copy sleeve returned four hundred basis points last quarter. The committee wants to know whether to add to it. Nobody in the room can answer that question, because nobody has separated the part that came from being long a market that went up, the part that came from filling behind the wallet you were following, and the part that came from having picked that wallet rather than another one.
Those three are different businesses with different capacities, different persistence and different reasons to stop working. Reporting them as one number is not a presentational shortcoming. It is a failure to know what you own.
The identity the report has to satisfy
Set up the decomposition so the terms sum to the realised sleeve return by construction, with a residual that you name rather than absorb. Four terms, in the order you should compute them.
- Beta. What a passive holding of the same instruments, in the same notional, over the same intervals, would have returned against your chosen market benchmark. This is the part you could have bought without a whale tracker.
- Selection. The return the tracked wallets themselves generated at their own fill prices over those intervals, in excess of the beta term. This is the value of having picked those wallets rather than the cohort at large.
- Timing. The difference between the wallets' fill prices and yours, on both legs, signed. This is the cost or benefit of arriving late.
- Residual. Fees, funding, borrow, and the trades you did not manage to replicate at all. Almost always negative, almost never reported.
Compute in that order and each term is defined against something already fixed, which stops the selection term from quietly absorbing everything you failed to model. Compute selection last, as a plug, and it will look excellent forever.
Timing is the only term you can measure exactly
Beta requires a benchmark choice and selection requires a cohort definition. Timing requires neither. It is arithmetic on four prices per round trip, and it is the term I would build first because it is the one nobody can argue with.

For each replicated trade, record the observed wallet fill price and timestamp, your fill price and timestamp, the notional, and the instrument. Timing on the entry leg is the wallet's price minus your price, times your quantity, signed for direction. Repeat on the exit. Sum across the sleeve and express it in basis points of average deployed capital.
Two things fall out of this immediately and both are worth having. The first is a realised latency distribution rather than an assumption, which is the input any future capacity work will need. The second is the sign. A desk that assumes timing is uniformly a cost is often wrong. On mean reverting instruments, arriving after the wallet's own market impact can occasionally fill you better, and knowing which regime you are in changes how aggressively you should chase.
Isolating beta so selection is not flattered by it
The tracked venue set spans DEX perpetuals on Hyperliquid, GMX, Drift and dYdX, and prediction market wallets on Polymarket and Opinion Trade. Those two legs need different benchmarks and combining them into one is the most common way a copy sleeve overstates its own selection.
The perpetuals leg has an obvious beta. A directional long in a rising market earns the market, and it earns it whether the wallet you followed was skilled or asleep. Benchmark it against passive exposure to the same underlying, in the same notional, held over exactly the same intervals. Interval matching matters more than instrument matching here, because a strategy that is only in the market during favourable stretches is claiming timing skill that must be evidenced, not assumed.
The prediction market leg has no comparable beta in the equity sense, and the honest treatment is to give it a benchmark of zero and let the entire return sit in selection. That is a more demanding standard, not a weaker one, and it will produce a selection term that is more volatile and more informative than anything the perpetuals leg gives you.
The cohort definition then has to be written down, because selection is measured against it. The Statistics view at capture reported a 44.3 percent average win rate across 26,687 tracked whales, while the Feed on the same day showed a 12.1 percent average across the 28 wallets active in that window. Those are two entirely different comparators, and a selection term computed against one is not comparable to a selection term computed against the other. Pick the one that matches your investable universe, which for most desks is the filtered subset closer to the 123 rows the Finder displayed than to the full census, and use the same one every period.
The fourth term nobody books
Replication shortfall is the trade the sleeve should have taken and did not. The wallet entered while your risk system was blocking new positions. The instrument was outside your approved list. The clip would have breached a concentration limit. The signal arrived at three in the morning and the desk is not staffed overnight.
These are invisible in a standard performance file because nothing was booked, and they are exactly the trades that most distort the picture. Copy strategies tend to miss trades non randomly, and the bias usually runs toward missing the fast violent ones, which are disproportionately either the best or the worst outcomes in the distribution.
The fix is a shadow log. Every observed signal that met your criteria gets a row, whether or not it was executed, with the reason for any non execution. Attribution then reports the sleeve's realised return alongside the shadow portfolio's return, and the gap between them is the operational cost of your own constraints, expressed in basis points. That figure belongs in the review, and it is frequently larger than the selection term everybody is arguing about.
What this lets you say when the sleeve is down
The reason to build the decomposition before you need it is that it converts a bad quarter from an argument into a diagnosis, and the four possible diagnoses call for four different responses.
Negative beta with positive selection and flat timing means the wallets were right and the market was against you, which is an exposure conversation about hedging the directional component, not a signal quality conversation. Negative selection with everything else intact means the wallets stopped working, and the response is cohort rotation with evidence rather than a defence of the original picks. Negative timing means the execution stack degraded, which is fixable and is nobody's investment thesis. And a large negative residual with the other three terms healthy means the strategy works and your operational implementation does not, which is the most actionable finding of the four and the one that never surfaces without the shadow log.
Without the decomposition, all four of those quarters look identical in the report, and the discussion defaults to whether people still believe in the idea. With it, the discussion is about a specific term, its size, and what changes it. That is the difference between a sleeve that survives a bad quarter and one that gets cut on sentiment.