The question is well posed and the test is worth running. If insider buying at the sector level turns before sell-side estimate revisions do, that is a genuine edge in sequencing and it is measurable. If revisions turn first, the insider series is confirmation rather than anticipation and should be weighted accordingly in whatever process consumes it. What decides the answer, in my experience, is almost never the estimator. It is four alignment decisions taken before the estimator runs.
The two series and where each one comes from
One leg is on the screen. The Insider Alpha Sector Flow panel publishes a rolling seven day window of insider buy value and transaction count across ten sector buckets. On the capture in front of me those read Industrials 88.92 M USD over 26 transactions, Consumer Cyclical 16.38 M USD over 11, Utilities 7.13 M USD over 6, Healthcare 4.26 M USD over 20, Basic Materials 3.02 M USD over 11, Financial Services 2.53 M USD over 52, Technology 979.40 K USD over 15, Consumer Defensive 432.68 K USD over 3, Energy 263.33 K USD over 5 and Communication Services 145.17 K USD over 7.
The other leg is not. There is no sell-side estimate revision feed inside this module, and I am not going to imply there is one. Revision breadth, however you define it, is a series you bring, from your own estimate database or a vendor. That has a practical consequence for the whole exercise: the two series will arrive on different taxonomies, different frequencies and different revision conventions, and every one of those has to be reconciled by hand.
Aligning taxonomies before aligning dates
Ten buckets appeared on the panel. The standard sector classifications carry eleven, and the row that was absent from this capture was real estate. Whether that is a permanent property of the module's taxonomy or an artefact of a week with no qualifying filings is exactly the sort of thing to establish before you build, because the two cases require different handling: a missing category is a mapping problem, an empty category is a data problem.
If your revision breadth series is computed on a different classification, the mapping you write is doing more work than the regression will. Sector definitions disagree most at the boundaries that matter here, with payment processors, consumer platforms and integrated energy names moving between buckets under different schemes. A lead-lag estimate that is sensitive to whether one large constituent sits in technology or in financials is not an estimate of anything.
The defensible approach is to map at the security level rather than the sector level. Build both series from constituent-level data using one classification you control, and treat the panel's ten rows as a cross-check on your reconstruction rather than as the input. If you cannot do that, restrict the test to the buckets where the two taxonomies agree unambiguously and accept a smaller cross-section.

Making the two series comparable
The insider series is a dollar-weighted count of discrete events, and the capture above shows how violently skewed it is. The revision series is typically a bounded diffusion index, upgrades less downgrades over total, which lives between minus one and one and moves smoothly. Correlating one against the other in raw units is a way of measuring which sector had the largest cheque this week.
Four transformations, in this order. Convert the insider leg to something bounded, either a within-sector z-score computed on a long trailing window or a cross-sectional rank across the ten buckets each period. Ranks are the safer default here precisely because of the skew. Second, decide whether the numerator is dollars or transactions and hold that decision fixed, since as the capture shows they rank the sectors differently at the top. Third, put both on a common frequency, which in practice means weekly, because the panel's native window is seven days and forcing a daily series out of a rolling seven day aggregate creates artificial smoothness and artificial autocorrelation. Fourth, fix the publication lag on each leg.
That last one deserves care. Form 4 carries a two business day filing deadline, so the insider leg is nearly current but not quite, and the transaction date and the filing date are different fields. Your revision series has its own vendor delivery lag and its own restatement behaviour. Both legs must be constructed point-in-time, using only what was visible on the date each observation is stamped with. A lead-lag result produced from a database that has been backfilled is a measure of the backfill.
The estimate, and what will make it lie
The mechanics are a cross-correlation of the two transformed series at weekly lags, per sector, pooled with sector fixed effects. That part is an afternoon. The interpretation is where the discipline is needed, and there are four specific ways this test produces a confident wrong answer.
- The effective sample is tiny. Ten sectors and a few years of weekly observations is not thousands of independent data points, it is a small number of sector-cycles observed ten times over with heavy cross-sectional correlation. Report the effective sample size, not the row count.
- Overlapping windows inflate significance. A rolling seven day aggregate sampled weekly is fine, sampled daily is not, and any smoothing applied to the revision leg compounds it. Use a block bootstrap for the standard errors rather than anything that assumes independence.
- Common drivers look like leads. Both series respond to the same macro shocks. If insider buying and revision breadth both fall when credit spreads widen, whichever series is measured with less lag will appear to lead. Include a control for the common factor, or at minimum report the result with and without one.
- The winner will not be stable. Run the test on two disjoint halves of the sample before reporting anything. If the lead flips sign across halves, the honest finding is that there is no reliable lead, which is a result worth having.
A placebo is cheap and worth building. Shuffle the sector labels on one leg and rerun. Whatever lead-lag magnitude the shuffled version produces is your noise floor, and any real result needs to clear it by a margin you specify in advance.
Using them together while the race is unresolved
The framing as a race is useful for the research question and unhelpful for the process, because the decision a desk actually faces is not which series to follow but what to do when they disagree.
The construction that survives review is a gate rather than a signal. Require both legs to point the same way before a sector tilt is expressed at full size, and express a reduced tilt when only one does. That is agnostic about which leads, it degrades gracefully if the lead-lag estimate turns out to be noise, and it produces a record of every occasion the two disagreed, which is the dataset that eventually answers the original question better than the regression will.
Log the disagreements with the date, the direction of each leg, the tilt actually taken and the reason. After enough of them you will have a point-in-time sample of exactly the situations the research was about, collected without look-ahead, which is a better foundation than any backfilled study you could run today.