MRE composites three named engines. AMS combines eleven breadth and on-chain metrics, MMS spans six timeframes of momentum, GLS tracks eight central bank balance sheets. The product describes the result as multi-signal confluence, and confluence is a claim with a specific statistical content: that combining the legs reduces the variance of the resulting score by more than combining three copies of the same thing would. That claim is testable, it is cheap to test, and it is the first thing a model risk reviewer should ask you for.
What confluence is claiming, in one line of algebra
Take three standardised leg scores with unit variance and pairwise correlation rho. The variance of their equal-weighted average is (1 + 2 rho) / 3. Read off the cases. At rho of zero the variance is 0.33, which is the textbook result that averaging three independent estimates behaves like a sample of three. At rho of 0.5 it is 0.67, and the effective sample size, defined as one over that variance, has fallen to 1.5. At rho of 0.8 it is 0.87, and effective N is 1.15. Three legs correlated at 0.8 give you roughly one leg and a rounding error.
This is the number to lead any internal write-up with, because it converts an argument about philosophy into an argument about a single measured quantity. Nobody has to agree about whether liquidity conditions and altcoin breadth are conceptually different things. You measure the correlation on your own history of the scores, you compute effective N, and you either have a three-leg model or you have a one-leg model wearing a three-leg label.

Three correlation matrices, not one
Correlation on the raw score levels is the matrix everyone computes and the least informative of the three. Level correlation is inflated by common trend, common seasonality, and the simple fact that all three legs are computed off market data that shares a regime. It will overstate dependence in calm periods and understate it in the periods you care about.
Compute the matrix on first differences of the scores as well. Innovation correlation asks whether the legs move together, which is closer to the question of whether they carry separate information. Then compute Spearman rank correlation alongside Pearson, because these are bounded scores that are unlikely to be jointly normal and a couple of shared extreme days can carry a Pearson estimate on their own.
Then compute the one that actually matches how the product is used: co-firing. The setup table emits discrete events, not continuous scores. Build the binary series of whether each leg was in its extreme band at each supported timeframe close, and measure the phi coefficient or the Jaccard overlap between those series within a tolerance window of a few bars. Two legs can be weakly correlated in levels and fire together ninety percent of the time, because both are thresholded functions of the same tail behaviour. Co-firing is the dependence that eats your confluence.
The conditional hit-rate test that decides it
Correlation tells you whether the legs are the same. It does not tell you whether the second and third legs earn their place. For that, partition the archived setups by which legs were in agreement and compare outcomes.
The partition is straightforward. For every archived setup, record the state of each leg at signal time, then bucket into one-leg, two-leg and three-leg agreement, keeping the direction fixed. Compute the realised hit rate and the mean outcome per bucket, net of costs, on a fixed horizon defined before you look. If confluence is real, the buckets order themselves and the gaps are larger than the sampling error. If they do not order themselves, the extra legs are filtering rather than confirming, and filtering has a very different capacity and turnover profile.
Put a standard error next to every one of those numbers or the exercise is theatre. At a hit rate near 0.5, the standard error is roughly 0.5 divided by the square root of the bucket count. With four hundred setups in a bucket that is 2.5 percentage points, so a three point lift from adding a leg is inside noise. With one hundred setups it is five points, and almost nothing you observe will be significant. The archive counter on the dashboard read 3,021 setups at capture, which sounds generous until you split it three ways by agreement, two ways by side, and again by timeframe.
The slow leg problem, and why low correlation can be worthless
GLS tracks eight central bank balance sheets. Balance sheet data publishes on a weekly to monthly reporting cadence. On a fifteen minute setup, which is what the entire visible feed in the capture above consists of, that leg is very close to a constant.
A near constant leg has two properties that look contradictory and are not. It will show low measured correlation with the fast legs, because it barely varies and correlation is a statement about co-variation. It will also contribute almost nothing to the cross-sectional ordering of signals, because a term that is the same for every signal this week cannot separate this week's signals from each other. Low correlation is being read as evidence of independence when it is evidence of low variance. Those are different findings with opposite implications.
The test that separates them is a perturbation. Re-score the archive with the slow leg held at its period mean, then again with it randomly permuted across periods, and compare bucket hit rates against the true configuration. If neither perturbation moves the outcome distribution meaningfully, the leg is not contributing at that horizon regardless of what the correlation matrix says. It may still be contributing at a weekly or monthly horizon, which is a separate test on a separate signal population, and worth running before anyone deletes anything.
Documenting it so it survives a review
The artefact that makes this defensible is short and it needs four things in it. The three correlation matrices, with the sample window and the treatment of overlapping observations stated explicitly. The implied effective N from the measured rho, so the reviewer sees the confluence claim as a number rather than an adjective. The conditional hit-rate table with confidence intervals per bucket and the bucket counts visible. And the perturbation results for any leg whose native update frequency is slower than the signal horizon.
Set a re-test cadence and record it in the same document, because correlation between legs is not stable. Legs built on breadth and legs built on momentum decouple in trending markets and converge hard in liquidations, which is precisely when a book relying on confluence is carrying its largest position. A number measured once in a calm sample and never revisited is the kind of finding that reads fine in a memo and fails in the month you need it.
One honest caveat belongs in the document too. Everything above is measured on the vendor's archive of its own signals, which is a population selected by the vendor's own thresholds. Independence measured inside that selection does not generalise to the legs as raw scores, and a reviewer who knows the difference will ask. The answer is to state the population you tested on and not to claim more than the sample supports.