Any statistic computed across the full population of quarterly filers is dominated by managers who are not expressing views. Wealth advisories running model portfolios, index complexes, insurance equity sleeves and pension plans holding external wrappers all file the same form as a concentrated equity fund, and there are vastly more of them. Aggregate without separating them and the resulting institutional-flow series mostly measures asset gathering and rebalancing, which is a real thing but not the thing anyone thinks they are looking at.
The instinct is to keep a list of interesting managers. That works for a while and then quietly rots, because the list reflects who was interesting when somebody last maintained it, and because a hand-curated list cannot be defended in a review beyond saying that it seemed reasonable. A computed taxonomy can be defended, reproduced, and back-dated, which are the three properties that matter when a flow-based input is questioned.
Five types, defined by what they do rather than what they call themselves
Regulatory categories are close to useless here, because a registered adviser can be an index shop or a concentrated activist and the registration says nothing. Define the types by portfolio behaviour instead.
- Index complex. Very large position counts, weights that track capitalisation, turnover close to the rebalancing minimum. Their filings are a mechanical restatement of an index.
- Wealth advisory. Large position counts, heavy use of fund wrappers, low concentration, turnover driven by client flows and model changes rather than views.
- Asset owner, meaning pension plans and endowments. Moderate position counts, low turnover, a book that is often mostly wrappers because the real risk is taken by external managers who file separately.
- Insurance. An equity sleeve that behaves like an index book, sitting under a balance sheet whose actual risk is in instruments that never appear on the form at all.
- Discretionary manager. Concentrated, higher turnover, dispersion in position sizing, options present, and a weight vector that has no particular relationship to any index.
Only the last type produces filings where the position is the point. Everything else produces filings where the position is a consequence of a mandate. The taxonomy exists to make that distinction machine-readable.

The three measures that carry the assignment
Three quantities do nearly all of the separation. Each is computed per filer per quarter from data already in the filings, which is what makes the scheme reproducible.
Turnover. Take the absolute share change for every line between consecutive quarters, price each at the later quarter-end price implied by the filing itself, sum, and divide by average reported portfolio value across the two dates. This is a coarse proxy and it understates real activity badly, because round trips inside the quarter are invisible. That is fine. It does not need to be accurate, it needs to be monotonic across manager types, and it is. Index complexes and asset owners sit near the floor. Discretionary managers sit well above it.
Concentration. Top ten weight is the readable version and a Herfindahl index on position weights is the version that behaves better in the tails. Compute both. A book where the top ten lines are a third or more of reported value is making decisions. A book where the top ten lines are a low single-digit percentage is holding a market.
Index overlap. This is the measure that does the most work and the one most schemes leave out. Take a broad capitalisation-weighted basket at the same date, restrict to names the filer holds, and measure how closely the filer's weight vector tracks the basket weights. High overlap with high position count is an index or near-index book regardless of what the entity is called. Low overlap is a manager choosing.
Two secondary flags refine the edges. Wrapper share, meaning the fraction of reported value in fund and trust lines, separates asset owners and wealth advisories from direct managers. Option presence, meaning any row carrying a put or call designation, is close to a positive marker for discretionary management, since the other types almost never carry them.
Thresholds, boundaries, and an unclassified bucket that stays honest
The temptation is to set cut points once and treat them as physics. Resist it. Set them as percentiles of the filer population in the same quarter, so the scheme adapts as the population changes and so a shift in the whole industry does not silently reclassify half your universe.
The assignment logic should be ordered rather than scored, because ordered rules are explainable and a composite score is not. Test for index overlap first, since a high-overlap high-breadth book is an index book no matter what its turnover looks like. Then test wrapper share, which pulls out asset owners and most advisories. Then test concentration and turnover jointly to separate discretionary managers from what remains.
Keep an explicit unclassified bucket and watch its size. A scheme that classifies everything is a scheme that is guessing on the boundary cases, and the boundary cases are exactly where a wrong label does damage. A manager with high concentration and near-zero turnover might be a long-horizon concentrated investor or a holding company, and those deserve different treatment. Leaving them unclassified and reviewing the list once a quarter is cheaper than the alternative, which is discovering the misclassification through a position.
Classify with a lag and stamp the label with an effective date. Assigning today's label to a manager's filings from four years ago is a look-ahead that will flatter any historical analysis you run, and it is invisible in the output. If a firm changed character, the history should show the old label until the quarter the measures actually changed.
Using the taxonomy without laundering it into a signal
The taxonomy is a filter and a control variable. It is not itself an edge, and the failure mode is treating it as one.
As a filter, its main job is to make the denominator of every flow statistic mean something. Institutional ownership change computed across discretionary managers only is a different series from the all-filer version, and the difference is the point. As a control, it lets you say something specific in attribution. When a flow-informed position works, being able to state that the flow came from concentrated discretionary books rather than from advisory rebalancing is the difference between a repeatable process and a story assembled afterwards.
The crowding question is where it earns its keep. A name being accumulated by many discretionary managers with high concentration is crowded in a way that matters for your exit, because those holders share your reasons and will reach for the same door. The same name being accumulated by advisories is a flow with very different persistence. A single ownership-change number cannot separate those two situations and a typed one can.
The reason to build this rather than buy a label is the same reason to build any classification you intend to defend. In Insider Alpha the institutional filings are presented as cross-reference material for the Form 4 feed, and a cross-reference is only as good as your ability to say which side of it you are trusting and why. When a committee asks why a flow input was weighted the way it was, the answer that ends the discussion is a rule, three measurements, and the effective date the label was assigned.