Most research platforms carry a free cash flow field. Most firms use several of them at once without deciding to. One analyst computes operating cash flow minus capital expenditure, a vendor feed applies its own capex definition that includes capitalised software, a sector team on hardware nets out finance lease payments, and the risk system pulls whichever field the warehouse happened to expose. Nothing breaks visibly. The screens run, the attribution report prints, and the numbers are wrong in a way that is correlated with sector, filing regime and vendor coverage.
This is a data policy problem wearing a valuation costume, and it is worth treating as one, because the fix is a procedure rather than an insight.
How a mixed definition surfaces as a screening artefact
The first symptom is a screen that keeps surfacing the same kind of company for no thesis you can articulate. Rank a universe on free cash flow yield where definitions vary by source and the top decile fills with names whose particular definition is the most generous one. Capitalised development spending is the usual culprit. A company that capitalises heavily shows lower capex under a definition that counts only tangible additions, so its yield looks better than an identical competitor that expenses the same activity. You have built a screen that ranks accounting policy and reads as though it ranks cash generation.
The second symptom is worse because it appears in client reporting. A value factor built on a mixed field loads partly on the vendor. Coverage is not uniform across market cap or geography, so the field switches source as you move down the universe, and the factor picks up the switch as if it were a signal. When performance is decomposed, the exposure line says value and the underlying reality says data provider.
The third symptom is time series discontinuity within a single company. Nothing about the business changed, but the free cash flow series has a step in it because a definition changed at the source, or an accounting standard moved part of a payment from one section of the cash flow statement to another. Any momentum or growth measure computed across that step is measuring the restatement.
The export is the boundary, so stamp it
The point where this gets locked in is more mundane than most people expect. It is the moment data leaves a screen and enters a spreadsheet.

The Company Valuation Engine scores a large universe on its own composite method, 4,420 companies on the board at capture, and internally that method is applied the same way to every row. The comparability guarantee ends at the export. Once three analysts have each pulled a file on different days with different filters, you have three universes and one filename convention, and reconciling them later costs more than doing this properly would have.
The restatement procedure
The work itself is unglamorous and finite. Six steps.
- Choose the canonical definition and write it in words. Not a formula reference, words: operating cash flow as reported, less total capital expenditure including capitalised software and capitalised development, less finance lease principal repayments. Whatever you pick, the test is that two analysts reading the sentence compute the same number from the same filing.
- Write the recipe at field level. Every input maps to a specific line of the cash flow statement with a stated fallback when the line is absent. Fallbacks are where silent divergence lives, so each one needs its own note.
- Back-fill from primary filings rather than from a vendor's history. Vendor history is often itself restated forward under a current definition, which will hide exactly the breaks you are trying to find.
- Store as-reported and restated side by side, never restated alone. When a portfolio manager asks why a number differs from a terminal, the answer has to be retrievable in seconds.
- Stamp every row with a definition version key. When the definition changes, mint a new version rather than editing history in place, and keep the old one queryable.
- Reconcile a sample by hand. Twenty companies across sectors and filing regimes, computed manually from the filings, checked against the pipeline. The reconciliation failures are the specification, and they will surface at least two cases the written definition did not cover.
The breaks you will hit and how to record them
A decade of history contains discontinuities that no definition can smooth away. They should be recorded as flags on the affected periods rather than quietly interpolated.
- Accounting standard changes. When lease treatment changed, operating cash flow rose for filers with large lease books because part of the payment moved into financing. That is a presentational step, not an improvement in cash generation.
- Reporting regime differences within one universe. Interest paid may sit in operating or financing depending on the standard the filer uses, which alone can move a free cash flow figure materially.
- Company restatements. The as-reported series and the restated series both exist and both are correct for different questions. A backtest must use what was knowable at the time.
- Large acquisitions and disposals. The entity changed. Growth rates spanning the transaction describe a merger, not performance.
- Fiscal year changes and reporting currency changes, which produce stub periods and translation effects that look like volatility.
Flagging these does two jobs. It stops a research process treating an accounting event as a signal, and it gives you a defensible answer when someone points at a chart and asks what happened in that year.
Attribution stops lying when the denominator holds still
The payoff is not a prettier database, it is that performance decomposition becomes truthful. With one definition applied across a decade and across the universe, a value tilt built on free cash flow yield is measuring one economic quantity. The exposure line means what it says. When the tilt underperforms, the conversation is about whether the exposure was right rather than about whose field the number came from.
It also makes the screen falsifiable. A ranked list built on a stable definition can be checked against realised cash generation over the following years, and the check either supports the definition or exposes it. A ranked list built on a mixed field cannot be checked at all, because a poor result is always attributable to data rather than to method, which is a comfortable position and a useless one.
Scoping the project so it finishes
Restating a decade for a broad universe from primary filings is expensive and there is no version of this where it is quick. Two ways to keep it finite. Scope by use: restate only the names and periods that feed a live screen or a reported factor, and leave the rest on as-reported with a clear label. And scope by depth: five years of clean, reconciled history on the definition you actually use beats ten years of history that nobody has hand-checked.
There is also a case for not doing this at all. If free cash flow is one input among many in a discretionary process, and nothing in client reporting depends on it, the honest answer may be to label the field as indicative, ban it from any ranked output, and spend the analyst months elsewhere. What should not survive is the middle position, where a mixed field is treated as a measurement and quoted in documents that someone will read back to you in a review.