The uncomfortable conversation goes like this. You reduced gross exposure in March. The drawdown you were positioning for did not arrive, the benchmark ran, and in February of the following year a client, a consultant or your own compliance function asks you to explain the decision. You open the Macroeconomic Risk Scorecard, and it shows you today: a combined M7 score of 29, a risk label of LOW, a regime of SLOWDOWN. None of that is what you were looking at in March, and you have no way to get back there.
This is a records problem, not an analysis problem, and it is entirely preventable at a cost of about four minutes per observation. What follows is the record I would keep and, more importantly, the reasoning about which fields are load bearing when somebody is genuinely trying to work out whether you had a process or a hunch.
The revised series will contradict your own memory
Two separate things work against reconstruction after the fact.
The first is data revision. Labour and output series from the BLS and the BEA get restated, sometimes materially and sometimes years later. A model that reads those series will produce a different value today for a date in the past than it produced on that date. So even a perfectly faithful re-run of the model gives you a number you never actually saw, and if the revision went against you, the re-run will make your March decision look worse than the evidence in front of you at the time.
The second is that a dashboard is a live view, not an archive. It is built to answer what is true now, which is the right design goal. It is not an obligation on the platform to preserve your decision context, and treating it as though it were is how desks end up with no record at all.

The implication is blunt. If the reading was not captured on the day, it did not happen as far as any later review is concerned, and no amount of confident recollection substitutes.
The nine fields
One row per observation, appended to a file that nobody edits. Nine fields, of which the last three are the ones that separate a real audit trail from a screenshot folder.
| Field | Why it is in the record |
|---|---|
| Observation timestamp | Establishes what was knowable. Everything else is meaningless without it. |
| Combined M7 score | The headline the decision will be remembered by, whether or not it drove anything. |
| Risk label and regime label | The categorical state, which is what a non-specialist reviewer will actually engage with. |
| Health grade | A second summary on a different scale, useful because it sometimes disagrees with the score. |
| Count of models at or above 60 | The agreement statistic. Without it you cannot later distinguish a broad signal from one loud model. |
| The seven individual model readings | The only field that supports genuine attribution. |
| Which pre-written rule was triggered | Names the rule by its identifier in the policy document, not by description. |
| What was not elevated | The inputs you checked that were quiet. This is the credibility field. |
| The pre-registered falsifier | What would have to read differently for you to reverse the decision. |
If the module surfaces a historical accuracy figure per model, capture that alongside the readings, because the reviewer's obvious follow-up question is whether the model you leaned on had any track record and you want the answer already in the file rather than assembled defensively afterwards.
Attribution is about which model, not which number
"The macro score was elevated" is not attribution. It is a restatement of the decision in slightly more formal language, and an experienced reviewer will treat it as such.
Real attribution names the component. It says the decision was authorised because a specified number of models crossed a specified threshold, and identifies which ones. That is a checkable claim. It lets somebody ask whether those particular models have a history of crossing early, whether they share inputs with each other, and whether the same combination has fired before without an event following.
The distinction matters most in the awkward case, which is a de-risking driven by one model. That is not automatically wrong. A single model can be reading something the others structurally cannot. But it is a materially weaker evidence state than three models agreeing, and a record that shows a count of one at or above 60 forces you to have made that argument at the time rather than discovering it under questioning.
The field I would fight hardest to keep is "what was not elevated". A record that only lists the alarming readings looks exactly like a record assembled to justify a conclusion, because that is what selective evidence looks like from the outside. A record that says credit spreads were unremarkable and the policy path had not changed, and we cut anyway for these two specific reasons, is a record made by someone who was weighing rather than confirming. Reviewers can tell the difference and it is the single cheapest way to buy credibility.
What this record wins, and what it does not
Be clear about the boundary, because overselling the audit trail is its own failure. This record wins the process question completely. Was there a rule, was it written before the observation, was the observation captured, did the decision follow the rule, was contrary evidence considered. Those are the questions a consultant or a compliance reviewer is actually mandated to ask, and a nine field row with a timestamp answers all of them in a way that no reconstruction can.
It does not win the outcome question, and you should say so before anyone else does. The record cannot show that cutting was correct, because a probability estimate being followed by no event is not evidence the estimate was wrong. Over any small number of decisions, a well-run macro overlay and a badly-run one look similar, and the honest framing in the room is that a single call cannot be evaluated on its outcome at all.
What the record does do over time is make the evaluation possible. Twenty or thirty dated rows, each with the readings, the triggered rule and the pre-registered falsifier, is a dataset about your own process. It will tell you which rules fire often and change nothing, whether your falsifiers ever actually get checked, and whether the count of models at or above 60 has any relationship to the decisions you went on to regret. None of that is available to a desk that keeps its macro reasoning in meeting notes, and all of it comes from four minutes a month of typing numbers into a file.