A composite score is a compression. It takes several measurements with different units, different noise properties and different degrees of relevance to your particular mandate, and it returns one number between zero and a hundred. That compression is genuinely useful for ranking a long feed quickly. It is also lossy in ways that matter, and the loss is invisible unless you know what went in.
Political Alpha documents its per-trade signal score as weighting four things, the politician's history, committee fit, the size of the trade, and its timing. At capture the dashboard's Avg Signal tile read 47.3, and the cluster cards carried their own scores, with a Microsoft cluster spanning three politicians over a nine day window scoring 45. Those numbers are only interpretable if you know what the components are doing, so this is an attempt to open the box and then to talk about what a desk should do with it.
The four inputs and what each is a proxy for
Take them in order of how much I trust them, which is not the order they are usually listed in.
Timing is the most objective input in the set and the one with the least ambiguity. It is the gap between the transaction date and the disclosure date, which the feed exposes directly through its Gap chips running from Same Day through to Late over 45 days. The average across the module read 32.5 days at capture. This is a clean, mechanically measured, non-judgmental quantity, and its relevance is obvious. Information decays, and the module quantifies that decay on its own performance tab, showing an average price movement of minus 1.88 percent between trade and disclosure.
Size is objective in what it measures and unreliable in what it means. Disclosures carry bands rather than amounts, so any size input is either an ordinal treatment of the band or a numeric conversion using an assumed point within it. The module's own average trade figure of $40K is described as a midpoint convention, which tells you the assumption being made. Band is a real signal in that a large band is a bigger commitment relative to the filer's own history. It is a weak signal in that the bands are wide and the smallest band is dominated by administrative activity.
Committee fit is the most interesting input and the most contestable. The module has a Committee Correlation view in its heatmaps section, which is how you inspect this dimension directly. The premise is that a transaction in an industry a member has oversight of is different to a transaction outside it. That premise is reasonable. The measurement is hard, because committee jurisdictions do not map cleanly onto sector classifications, memberships change, and the same overlap can reflect ordinary familiarity rather than anything else.
Politician history is the input where I would want the most detail before leaning on it. The leaderboard exposes the per-member numbers this would draw on, showing trade counts, a 30 day alpha column, a win rate and a compliance score. At capture one House member's row showed 9,596 trades, an alpha of minus 1.1 percent and a 44 percent win rate. Track record over a member with thousands of transactions is a statistically meaningful measurement. Over a member with eleven, it is not, and the score cannot tell you which case you are in.

Score dispersion matters more than score level
An average of 47.3 on a zero to hundred scale tells you the scale is roughly centred, and nothing else. The question that determines whether a composite is usable is dispersion, and specifically whether the top decile is meaningfully separated from the middle.
Pull the distribution before you use the score for anything. If most filings cluster between 40 and 55, then the difference between a 52 and a 48 is noise and any threshold you set in that range is arbitrary. If the tail is thin and well separated, a threshold is defensible. This is a ten minute exercise and it should precede any conversation about what cutoff to use, because a cutoff chosen without looking at the distribution is a number somebody made up in a meeting.
Do the same by component if you can reconstruct them. A composite where one input dominates the variance is effectively a single-factor score wearing a coat, and you should know that before you describe it to an allocator as multi-factor.
Rebuilding the composite against your own mandate
Here is the honest framing. You are not tuning the vendor's model. You are building your own composite from the same underlying fields, which the module exposes individually, and then deciding whether the vendor's default or yours ranks better for your purposes.
The mandate-specific reweighting is straightforward once you state what you actually need. A capacity-constrained book should upweight size and downweight everything else, since a small band trade in an illiquid name is unusable regardless of how interesting it is. A book with a slow implementation path should upweight timing severely, because a filing that is already three weeks old plus your own two week research cycle leaves nothing. A thematic or policy-oriented mandate should upweight committee fit and accept the measurement noise, because that dimension is the one connected to the story you are actually telling clients.
Whatever you land on, freeze it, version it, and record the date. A score whose weights drift as somebody tinkers is not a model, and the first question in a review will be whether the weights predate the results.
Validate against the module's own numbers, not against your hopes
The performance tab is where this gets uncomfortable and where the honest work happens. At capture, across 5,000 trades analysed, it showed an average 30 day alpha of minus 0.25 percent, an average 90 day alpha of minus 0.34 percent, and a 48.2 percent win rate.
Read those as the unconditional baseline for the cohort. They are not a verdict on the score, because the score exists precisely to select a subset that behaves differently to the average. But they set the bar the score has to clear, and they set it in a specific way. A scoring system on a population with a slightly negative unconditional mean has to be doing real selection work rather than mild reordering. If your own conditioning on high scores produces a distribution that overlaps heavily with that baseline, you have a ranking tool and not a signal, and the difference determines whether this belongs in a mandate or in a research folder.
Where a single number stops being safe
Two failure modes are worth stating plainly. The first is compensation between components. A composite lets a very high score on one dimension mask a disqualifying value on another, so a large trade by a member with a poor track record filed forty days late can score respectably. If any component has a level at which you would refuse the trade regardless, that is a gate, not a weight, and it belongs outside the composite.
The second is population instability. The dashboard tile counting active clusters read 1 at capture, and 24 trades for the week against 870 for the month. When the underlying volume swings that hard, a percentile-based score is being computed against a different population week to week, and a threshold that was selective in a busy month becomes indiscriminate in a quiet one. Fix your thresholds to absolute values or recalibrate them on a stated schedule, and write down which one you chose.