The Macroeconomic Risk Scorecard dashboard reads 29 out of 100 with a risk label of LOW, and the tile beside it says the regime is SLOWDOWN. If your sizing process consumes that 29 and nothing else, you have thrown away the more useful half of the reading. The header row says it plainly: 7 MODELS, COMBINED M7 SCORE. A summary of seven numbers is a different object from the seven numbers, and for a book that has to justify its gross exposure in a quarterly review, the difference is where the information lives.
This is not an argument about which recession model is best. It is an argument that the second moment of the model set is a usable input and that most processes throw it away at the point of ingestion, because the platform hands you a clean scalar and a scalar is easy to put in a spreadsheet.
Two books can print the same 29
Suppose all seven models land between 24 and 34. The composite is 29 and the cycle read is genuinely quiet. Seven methods built on different families of input, yield curve behaviour, credit spreads, labour and output series, policy settings, are all saying the same unexciting thing. Your uncertainty about the level is small because the evidence is not in conflict.
Now suppose four models sit near 10 and three sit near 60. The composite is still somewhere near 29. That book is not in a quiet market. A meaningful minority of the evidence is already elevated and the average is hiding it. The headline number is identical and the risk state is not.
Neither of those is a forecast, and I want to be careful about that. Dispersion does not tell you what happens next. It tells you how much weight the level deserves, which is a different and more modest claim, and it is the claim that actually maps onto sizing.
The agreement statistic is already on the tile row
The scorecard header carries a tile labelled MODELS at or above 60 alongside the combined score and the regime. That counter is the coarse version of what I am describing. It tells you how many of the seven have crossed into elevated territory regardless of where the average sits, which is precisely the piece the composite averages away. Zero of seven at a composite of 29 and three of seven at a composite of 29 are different states, and only one of them is visible in the headline.

If you want something finer than a counter, record all seven readings at each observation rather than the composite alone, and carry two numbers forward. The mean, which is close to what the combined score gives you, and the spread, whether you express that as the standard deviation across the seven or simply as the gap between the highest and lowest reading. The mean is the level. The spread is your confidence band around it. Both belong in the same monthly row of your macro log, and the log matters more than the sophistication of the statistic.
Turning the spread into a gross exposure multiplier
The mapping below is a policy, not a platform output. The module gives you the readings. The third column is a number you choose, write down in advance, and defend later. That distinction is the whole point of doing this in a process rather than in your head.
| State | What the model set is saying | Gross multiplier |
|---|---|---|
| Tight, low mean | Seven independent methods agree nothing is breaking | Full risk budget |
| Wide, low mean | The average is calm because disagreement is cancelling out | Trim, because the level is not trustworthy |
| Tight, high mean | Broad agreement the cycle is turning | Cut on the level, not on the uncertainty |
| Wide, high mean | Elevated and contested, the worst state for conviction | Smallest gross, longest decision horizon |
Two design notes on that table. First, the wide and low cell is the one that earns its keep, because it is the only state where a level-only process leaves you fully invested and a dispersion-aware process does not. Second, the trim should be modest. If your dispersion rule can move gross by more than about a fifth of the budget, you have built a macro timing strategy and given it a risk management label, and it will be attributed as timing whatever you call it.
Where dispersion scaling reliably fails
It fails in fast breaks. Dispersion is a lagging description of the evidence set, so in a shock that moves every input in the same week, the spread collapses and the mean jumps at the same time. You get the level signal and the agreement signal simultaneously, which means the dispersion overlay added nothing at the moment you most wanted it. Anyone selling you this as an early warning system is selling you something the construction does not support.
It also fails on the independence assumption. Seven models are not seven independent opinions in the statistical sense. If several of them consume overlapping input families, and any macro model set built on the same fifty plus indicators from FRED, the BLS, the BEA and the ECB will, then part of the agreement you observe is mechanical rather than evidential. Tight dispersion in that setting is weaker evidence than it looks. You cannot fully correct for this without knowing the full construction of each model, so the honest response is to keep the multipliers small rather than to pretend the correction has been made.
And it is bounded. A composite near the floor or the ceiling of the zero to one hundred range compresses the possible spread mechanically. At a reading of 29 there is room for wide disagreement. At a reading of 6 there is not, and a tight spread there tells you almost nothing.
Writing it into the risk budget so it survives a review
Three things make this defensible rather than discretionary. The first is that the rule is written before the observation. A dispersion band that gets discovered in the month you wanted to reduce gross anyway is not a process, and an investment committee that has seen a few of these will spot it.
The second is the dated snapshot. Record the observation date, the combined score, the regime label, the count at or above 60, and all seven readings if you are capturing them. Macro data revises. The reading you acted on and the reading the series shows a year later are frequently not the same number, and without the snapshot you will lose the argument about what you knew, not because you were wrong but because you cannot show it.
The third is attribution honesty. A dispersion overlay scales gross. It does not pick names. So it will show up in attribution as timing, it will look like noise over any short window, and it needs to be evaluated over a horizon long enough to contain several regime transitions, which is longer than most review cycles. Say that at the outset. The alternative is explaining, in a quarter when the overlay cost you thirty basis points against a benchmark that stayed fully invested, why a risk control you never described in advance was reducing exposure on a number nobody in the room had seen.