An event contract book produces the least informative P&L series in the business. Every position resolves to zero or one hundred, so the return distribution is a pile of total wins and total losses, and the path in between is noise generated by an order book that may have printed four times all quarter. Hand that series to an allocator and there is nothing in it to underwrite. They cannot tell whether you are good at forecasting, good at buying, or were simply short a lot of things that did not happen.
The fix is to stop reporting the P&L as the primary evidence and start reporting a decomposition alongside it. The apparatus for this already exists, it is old, and it is not difficult. What it requires is that you record your fair value estimate at the moment of entry, every time, without exception, because the whole attribution collapses if that number is reconstructed afterwards.
What a good Brier score can be hiding
The Brier score is the mean squared error of your probability forecasts against the realised outcomes, so lower is better and zero is perfect. Reported alone it is close to useless, and it is useless in a direction that flatters you.
Consider a book made entirely of contracts priced near the far ends, of the kind that populate the top of the markets table when it is sorted by volume, with Yes columns reading 0.9 percent and 1.2 percent. Sell those, be right almost every time, and your squared errors are tiny. The Brier score will be excellent and you will have demonstrated nothing at all, because the outcomes were nearly deterministic before you arrived. That is the environment's contribution, not yours, and any attribution that does not strip it out is a marketing document.

The three-way split, and which term is yours
The standard decomposition breaks a Brier score into three parts. There is an uncertainty term, which is the variance of the outcome base rate over your sample and belongs entirely to the environment. There is a resolution term, sometimes called discrimination, which measures how far the realised frequencies in your forecast buckets sit from the overall base rate. And there is a reliability term, which is your calibration error, the squared gap between what you said and what happened in each bucket.
Run it as an illustration, with numbers you should replace with your own. Take four hundred settled contracts where thirty percent resolved Yes. The uncertainty term is 0.3 times 0.7, or 0.21, which is also the Brier score of a manager who forecasts the base rate on everything and does no work. Suppose your book scores 0.17. Decomposed, that might be uncertainty 0.21, minus resolution 0.05, plus reliability 0.01.
Now the three lines say three different things. The 0.21 was handed to you. The 0.05 of resolution is the only line that is evidence of forecasting skill, because it says your buckets separated outcomes. The 0.01 of reliability is an error, and importantly it is a cheap error, because miscalibration is repairable by a monotone transform of your stated probabilities without any new information. If reliability is large, you do not have a research problem, you have a mapping problem, and you can fix it this quarter.
Express the headline as a skill score against the base-rate benchmark, which here is 0.04 out of 0.21, or about nineteen percent. Then run the same score against the harder benchmark, which is the market price at the moment you traded. That comparison is the one that matters for a trading book, because the market price was available to you for free. If your Brier is not better than the Brier of the prices you paid, then whatever produced the P&L, it was not forecasting.
Count events, not contracts
Every statistic above assumes independent observations, and event contract books systematically violate that assumption in a way that inflates apparent sample size by an order of magnitude.
The trending page in this module makes the point better than any explanation. The most liquid board shows ten rows, each carrying between $3.8M and $4.7M of liquidity, and all ten rows are the same question about the 2028 Democratic presidential nominee, listed once per candidate. A book holding all ten of those legs holds ten contracts and one event. When the nominee is chosen, all ten resolve simultaneously and their outcomes are perfectly determined by each other. Reporting that as ten observations is not a rounding error, it is a tenfold overstatement of the evidence.
So the attribution page reports two counts. Contracts settled, and effective independent events after clustering by underlying resolution driver. Cluster conservatively. Contracts that resolve from the same data release, the same match, the same election or the same regulatory decision are one event. Then quote your confidence intervals on the effective count, and if the effective count is under a hundred, quote intervals rather than point estimates and say so plainly. An attribution with sixty effective events cannot distinguish a good manager from a lucky one, and saying that yourself is considerably better than having it said to you.
The execution line the Brier score cannot see
Forecast quality and trading quality are different skills and they are routinely conflated. A well calibrated book that pays through the offer every time will lose money. A mediocre forecaster who only ever gets filled on resting orders can grind out a return. The Brier score sees none of this, because it never looks at the price you paid.
Add a dollar decomposition next to the probability one. At entry, record your fair value and the price you transacted, and the expected edge is size times the difference, summed across the book. That is the P&L you told yourself you were buying. Then set realised P&L beside it. The residual is resolution variance plus model error, and over enough independent events it should shrink toward zero if the model is honest.
Break out fees and financing separately rather than netting them into the residual, because they are the one component that is fully knowable in advance. And add a slippage line comparing the price you intended against the price you got, which on thin binaries is often larger than the entire forecast edge. On this asset class the Liquidity column can read a few hundred dollars on a market doing hundreds of thousands in daily volume, and a book that ignores that gap will show a forecast edge and a realised loss with no line item connecting them.
The page that survives a manager review
One page, per category and then aggregated, per period. Contracts settled and effective independent events. Base rate and the uncertainty term. Brier score, reliability, resolution. Brier of the market prices at entry. Skill score against the base rate and against the market. Expected dollar edge at entry, realised P&L, fees and financing, slippage, and the residual explicitly labelled as resolution variance rather than buried.
The reason to build this before you need it is that the conversation it enables runs in one direction only. A manager who arrives at a review with a decomposition can say that the quarter's loss sat entirely in the residual, that resolution held at 0.05, and that the process is intact. A manager who arrives with a P&L chart is asking to be believed, and after a bad quarter on an instrument that pays zero or one hundred, belief is the one thing nobody in the room has left.