At capture the MRE dashboard was paging through an archive of 3,021 setups, showing 100 of them at a time. The tickers-covered tile read 4 over the last thirty days and the setup counter was averaging 30.1 per day over the previous week. Those four numbers are the entire basis on which anyone can claim the engine has an edge, and the first job of due diligence is working out how many independent observations they represent. It is not 3,021.
What the counter counts
A row in that archive is one signal emission. It is not one trade, not one market event, and certainly not one independent draw from a distribution. Three distinct sources of dependence sit between the raw count and anything you can put a t-statistic on.
Cross-sectional dependence comes first. The covered instrument set in the capture is small and internally correlated. The visible rows are BTC/USD, ETH/USD and stablecoin dominance, and a fourth instrument sits behind the count. In crypto beta terms those are not four assets. A common factor drives most of the joint variation, and stablecoin dominance is close to the arithmetic complement of the rest of the complex rather than an independent instrument. If you are being generous you have between one and two independent cross-sectional units, not four.
Temporal clustering comes second. The top two rows in the capture are both BTC/USD, both 15m, both OVERBOUGHT and SHORT, both SCALP, both carrying an identical Read string, timestamped roughly sixteen hours apart. Those are two rows describing one persistent condition. The archive is full of that pattern by construction, because a condition that produces a signal on one bar frequently produces another on a nearby bar.
Timeframe nesting comes third. The engine runs multi-timeframe reversal detection, so a condition visible on a fifteen minute chart is often visible on the hourly view of the same instrument at the same moment. Counting both as observations double counts one event, and the double counting is not random, it concentrates in exactly the large moves that dominate any performance summary.

Collapsing rows into episodes
The unit you want is an episode, defined before you look at outcomes. A workable definition is this: same instrument, same side, within a window of k bars of the coarsest timeframe involved, collapses to one episode, taking the first emission as the episode timestamp. Choose k on the persistence of the underlying condition rather than on convenience, and state it. Then apply a second collapse across instruments: signals on correlated instruments inside the same window are one cross-sectional episode, weighted by however many names participated if you want to keep that information.
Run that on the archive and report both numbers, always together. Raw emissions of 3,021 and episodes of whatever you get. The ratio is itself the finding. If a 15m feed running at thirty emissions a day across four correlated names collapses at anything like the rate the visible rows suggest, you should expect to land in the hundreds of episodes, not the thousands. I am not going to guess the exact number for someone else's window choice, and neither should any document you sign.
The reason to do this before anything else is that every subsequent statistic inherits the choice. Hit rates computed on emissions inherit the clustering. Standard errors computed on emissions are too small by roughly the square root of the clustering factor, which means a result that looks significant at three standard errors on emissions can be comfortably inside noise on episodes.
The power arithmetic, in one line
The approximation worth memorising for a two-sided test at five percent with eighty percent power, comparing a hit rate against a fixed fifty percent baseline, is that the required sample is about 8 times p times one minus p, divided by the squared effect size. The 8 is the squared sum of the two critical values, 1.96 and 0.84, rounded. At a baseline of one half that collapses to roughly 2 divided by delta squared.
Read the consequences off that. To detect a five percentage point edge you need on the order of 800 independent observations. For ten points, roughly 200. Invert it and it is more useful still: the smallest edge an episode count of n can resolve is about the square root of 2 over n, so 400 episodes resolves down to about seven points and 300 episodes to about eight. That is the whole conversation. If your collapsed count lands in the low hundreds, the archive can settle whether the engine has a large edge and is structurally incapable of settling whether it has a small one. A small edge is not disproved by that. It is outside what this evidence base can speak to, and saying so is the defensible position.
The accrual rate does not rescue you as fast as it looks. Thirty emissions a day is roughly nine hundred a month, but they arrive on the same four correlated instruments, so they buy episodes at the collapsed rate rather than the headline rate, and they arrive within one market regime. Twelve months of accrual on the current instrument set adds sample but adds very little independence in the dimension that matters for a reversal strategy, which is how the engine behaves across different volatility and trend environments.
What happens when you slice it
The pressure on the sample gets much worse the moment anyone asks a sensible follow-up question, and they will. Does the edge hold on both sides. Does it hold on the fifteen minute timeframe as well as the slower ones. Does it hold on the scalp class specifically. Each of those is a legitimate question and each one multiplies the number of cells.
Suppose four timeframes, two sides and three setup classes, which is twenty four cells. Split 3,021 emissions evenly and you have 126 per cell before any clustering adjustment at all. The standard error on a hit rate near one half at 126 observations is about 4.5 percentage points, so a per-cell confidence interval spans roughly eighteen points end to end. On episodes rather than emissions, the interval is wider still. Cells will not be evenly populated either, and the visible feed suggests one timeframe dominates heavily, so the thin cells will be far thinner than the average implies.
There is a multiplicity problem stacked on top. Twenty four cells tested at five percent produces roughly one spuriously significant cell by chance even if the engine is worthless. Anyone presenting the best cell from a slice like that, without saying how many cells were examined, is presenting noise with a decimal point on it. Pre-register the slices you intend to test, correct for the number of tests, and report the cells you looked at and rejected alongside the ones you kept.
What this archive can honestly settle
It can settle direction and magnitude for a large effect on the pooled sample. It can settle whether the engine's emissions are, on a cost-adjusted basis, better or worse than a naive benchmark by a wide margin. It can characterise the shape of the outcome distribution, which is often more informative than the mean and needs less sample to see, particularly the tail behaviour that determines how a reversal book draws down.
It cannot settle modest per-cell differences, it cannot settle stability across regimes it has not lived through, and it cannot settle anything about instruments outside the covered set. Those are not criticisms of the module. They are properties of any archive built at thirty emissions a day on a handful of correlated names, and the honest move is to write the limits into the evaluation memo at the same time you write the headline number, rather than waiting to be asked about them in the meeting where the allocation is decided.