Every leaderboard invites the same question, and it is the right question. If a filer sits in the top decile this year, where do they sit next year? A rank that persists is a rank you can build a process on. A rank that does not is a description of the past dressed as a forecast. The test is straightforward to specify and easy to run badly, and most of the ways it goes wrong are decided before any data is loaded, in the definition of what is being ranked.
What the leaderboard rank actually is
Start with the mechanics of the board rather than the idea of it. The Top Insiders table on Insider Alpha ranks on a column headed 30d Buy $, and on the capture in front of me it ran from 100.00 M USD down through 98.71, 64.80, 54.24, 50.00, 41.35, 34.92 and 21.26 M USD across the visible rows. The header tile above it, labelled TOP INSIDER 7D, showed 589.61 M USD, so the summary tile and the table are measuring over different windows.
That is a flow ranking. It orders filers by how much capital they deployed in a window. It is not ordering them by realised return, and on this capture it could not have been, because the Win Rate column was blank on every visible row, as were Holdings $ and Top % Float.
This distinction decides your whole test. Persistence in a dollar flow rank is largely persistence in the capacity to deploy dollars, which for the entity filers that populate the top of the board is a property of their fund size and their mandate rather than of their insight. A sponsor that bought heavily last year will plausibly buy heavily next year for reasons that have nothing to do with whether the first purchase worked.
Why the naive test answers the wrong question
Take last year's top decile by dollars, look at this year's top decile by dollars, count the overlap, publish a transition matrix. You will get a number, and the number will probably look encouraging, and it will be nearly uninformative.
The reason is that a dollar rank is dominated by the size of the filer's balance sheet, which is the most persistent variable in the entire dataset. You will have measured the autocorrelation of fund size. If your investment committee reads that as evidence that top insiders repeat, they will start allocating attention to the largest buyers, which is exactly the population whose transactions are least likely to be discretionary open-market selection.
The fix is to decide, in writing, which of two hypotheses you are testing. Either you are testing whether a filer's realised forward returns persist, in which case the ranking variable has to be a return measure and the leaderboard dollar column is not it. Or you are testing whether the flow itself carries information, in which case persistence of the rank is beside the point and the test you actually want is whether the flow predicts returns at all, which is a different construction entirely.

Constructing the decile portfolios so the test can fail
Assume you have chosen the return-persistence hypothesis. The construction then needs five decisions made up front.
Formation universe. Set a minimum transaction count for inclusion and set it before you look at anything. The visible trade counts in the screenshot above are a seven day snapshot, but they are indicative of the shape you will find over a year: a mass of filers with one or two transactions and a thin tail with enough for an estimate. A filer with two transactions cannot be assigned to a skill decile in any meaningful way, and including them fills your extreme deciles with noise, which is the classic way to manufacture apparent mean reversion.
Point-in-time membership. The decile a filer belongs to must be computable from information available on the formation date. This sounds obvious and is routinely violated, because the natural way to build the panel is to pull the current filer list and then look backwards, which silently excludes everyone who stopped filing.
Survivorship. Filers exit. Officers retire, funds exit positions and drop below the ten percent threshold, companies delist. A persistence test run on filers who are still filing at the end of the sample will find persistence, because continuing to file is correlated with the underlying position having worked out.
Identity. The board displays names in inconsistent form, with some rows in full capitals and others in mixed case. Name strings are not stable identifiers. Key the panel on the filer's central index key from the filings themselves, and treat any name-based join as a source of both false merges and false splits.
The null. Report a Spearman rank correlation between adjacent years and a full transition matrix, but also report what the same statistics look like under a null built by shuffling filer labels within formation year. With a few hundred qualifying filers and a fat-tailed return distribution, the sampling variation is large enough that eyeballing a transition matrix is not a test.
The failure modes that manufacture persistence
Four of these will bite before anything else does.
- Same-name duplication. Two of the eight visible rows on this board pointed at the same ticker. If two filers in your top decile are expressing one situation, your decile has fewer independent bets than it appears to, and the correlation across years inherits that.
- Sector concentration. A decile that is three quarters one sector will persist for as long as that sector does. Attribute the decile return to sector before you interpret it, and report the sector composition of each decile alongside the transition matrix.
- Overlapping windows. Filers who buy repeatedly generate overlapping forward-return windows, so the observations inside a decile are not independent and the standard error you compute will be too small.
- Rank computed after the outcome. If the ranking variable uses any data from inside the evaluation window, the test is circular. This is easy to introduce accidentally when the ranking variable is a rolling measure.
What a negative result is worth
Run this properly and there is a real chance the answer is that the rank does not persist. That is a usable finding, not a wasted quarter, and it is worth saying in advance what you will do with it so that the finding does not get quietly reinterpreted.
A rank that does not persist can still be a perfectly good attention allocator. Knowing which filers deployed the most capital in the last thirty days is genuinely useful for deciding what to read, and that use survives a failed persistence test intact. What does not survive is any process that sizes positions on the strength of a filer's position in the ranking, or any marketing language that implies the board identifies durably skilled individuals.
Write both outcomes down before you run it, including the specific decisions each one changes. A persistence test whose negative branch has no consequences is not a test, it is a formality, and the review that eventually asks what evidence supported the process will find nothing behind it.