Most quality factors in production are a ranking on current return on invested capital with a leverage screen bolted on. They work in the sense that they produce a defensible tilt, and they fail in a predictable way: the level of returns is the most mean reverting thing about a company, so a factor that buys the level is systematically buying names near the top of their own cycle and selling names near the bottom of theirs.
The information a quality factor is supposed to capture is not how high returns are today. It is how long they stay high. That is a different estimand, it needs a different construction, and the construction is not difficult. What it needs is a decade of restated history per name and a willingness to accept a much noisier per-name estimate in exchange for a much more stable portfolio.
Why the level carries so little of the signal
Take the spread between return on invested capital and cost of capital rather than the raw return, since a fifteen percent return means different things in software and in utilities. Across any broad universe the cross-sectional distribution of that spread is wide at any point in time and narrows sharply as you extend the horizon. Competition is the mechanism, and it is the oldest result in the industrial organisation literature applied to accounting returns. High spreads attract entry, entry compresses them.
The practical consequence for a factor builder is that the top decile on current spread is populated by three distinct groups that a level ranking cannot separate. There are businesses with a structural moat whose spread will still be positive in ten years. There are cyclicals at the peak of their cycle whose spread is about to invert. And there are companies whose invested capital denominator is artificially small, usually because they expense the thing that is actually their asset, which describes most research-intensive and brand-intensive businesses. Rank on the level and you buy all three in roughly equal weight.
The estimation, which is one regression per group
The construction is an autoregression on the spread. For each company, take the annual spread series over as long a window as your restated data supports, ten years being the reasonable minimum, and estimate the coefficient on the lagged spread with a fade toward a long-run anchor. The coefficient is your persistence parameter. A value near one means the spread does not decay. A value near zero means this year tells you nothing about next year.
The half-life falls straight out of the coefficient. A persistence of 0.8 implies a half-life of about three years, meaning a spread of eight points is expected to be four points in three years and two points in six. A persistence of 0.6 implies a half-life closer to fourteen months. That single number is more useful to a portfolio manager than any level, because it is the assumption a discounted cash flow model makes implicitly and usually without stating.
Per-name estimation on ten annual observations is statistically thin, and pretending otherwise is how a quality factor becomes an overfitting exercise. The fix is hierarchical. Estimate the persistence coefficient at the sector or business-model group level where you have hundreds of company-years, then shrink each company's own estimate toward its group prior in proportion to how few observations and how much residual variance that company has. A name with fifteen clean years and a tight fit keeps most of its own estimate. A name with six years and a wide residual band ends up close to its group.

From a fade coefficient to a factor score
Three components, combined with fixed weights set once and documented, because a weighting scheme that gets revisited each quarter is a discretionary overlay wearing a factor label.
The first component is the shrunk persistence coefficient itself, ranked cross-sectionally within sector. The second is the current spread, included with a deliberately small weight, because you still want a positive spread rather than a durable negative one. The third is the precision of the persistence estimate, which enters as a penalty rather than a reward. Two names with the same estimated persistence of 0.8 should not score the same if one has a standard error twice the other's. Penalising imprecision is what stops the factor from loading on companies with short or messy histories, which is a real and reproducible failure mode.
Report the resulting score as a percentile within sector, never as a raw value, and keep the three components visible in the name-level record. When a position is questioned in review the answer needs to be that this name scored in the top decile on persistence with a tight estimate and a modest current spread, not that it scored 87.
Benchmarking against the sub-score already on the dashboard
Any platform that ships a fundamental sub-score gives you a free benchmark, and you should treat beating it as the bar rather than as a formality. The Company Valuation Engine carries a fundamental sub-score alongside technical and sentiment components, fused into one composite per company across a universe that read 4,420 companies with 4,432 fully scanned at the time of capture. That is a genuine breadth advantage over anything a small team restates by hand, and it is the reason to benchmark rather than dismiss.
Run three comparisons. Rank correlation between your persistence score and the vendor fundamental sub-score across the overlapping universe tells you whether you have built something different at all. If the correlation is above about 0.7 you have rebuilt the vendor's factor at considerable cost and should say so. Decile overlap, specifically the share of your top decile that also sits in the vendor's top decile, tells you where the difference concentrates. And a turnover comparison, since the entire claim of a persistence construction is that it changes its mind less often, which should show up as a materially lower annual name turnover and therefore lower implementation cost in basis points.
Where the two disagree hardest is the useful output. Names your factor likes and the sub-score does not are usually businesses with a modest current level and an unusually flat fade. Names the sub-score likes and yours does not are usually peak-cycle. Both lists are worth an analyst hour before any of them become positions.
The universes where the construction should be switched off
Return on invested capital is not defined in a useful way for banks and insurers, where the balance sheet is the product and leverage is the business model. Exclude them explicitly rather than letting a garbage estimate flow into a ranking.
Serial acquirers are the second exclusion, or at least a flagged category. Their invested capital jumps discontinuously with each deal, which makes the spread series a sequence of level shifts rather than a decaying process, and the autoregression will read those shifts as low persistence. That is a measurement artefact, not a finding about business quality. Either estimate on goodwill-adjusted capital consistently or carve them out and handle them by hand.
The third is the research-intensive and brand-intensive names where the capital base is understated because the asset is expensed. Their spreads look enormous and unusually stable, which the factor will read as elite quality. Capitalise development spend across the whole universe before estimating, or accept a known and directional tilt toward companies whose accounting flatters them.
On capacity and crowding, the honest position is that a low-turnover quality construction is capacity friendly relative to almost anything else on a desk, and that the crowding risk is not in the factor mechanics but in the names. A persistence-ranked top decile in a mature universe will contain a large number of holdings that every other quality manager also owns, and the correlation of that book to the rest of the quality complex is the exposure worth monitoring. Report it as a monthly number rather than discovering it in a drawdown.