On the Prediction Alpha markets tab there is a control labelled Cross Match with a threshold selector next to it, set to 80 percent in the capture I am working from, and a count of 79 beside it against 1,000 displayed rows. An Arb Only toggle carries the same 79. That threshold is the only thing standing between a question-matching model and a trade ticket, and on most desks nobody has written down why it is where it is.
The instinct is to treat it as a precision dial. Turn it up and fewer bad matches survive. That instinct is wrong in a specific and expensive way, and the reason is worth an afternoon of somebody's time before it costs you a position.
What the dial is actually ranking
A cross-venue matcher is comparing two questions written by two different venues and scoring how alike they are. Set the gate at 80 and you keep the pairs that score above 80. Set it at 95 and you keep fewer. That is a recall and precision tradeoff in the standard shape, and if the errors were distributed randomly across the similarity range, raising the gate would be a clean control.
They are not distributed randomly. The most dangerous false positive in this asset class is a pair of contracts from the same venue's own series, written from the same template, differing by one token. That pair scores near the top of the similarity distribution, not the bottom.

The failure modes string similarity cannot see
Look at the trending board on the same module for the cleanest demonstration. Three separate rows in the 24-hour volume leaderboard read Natus Vincere vs Fnatic Game 1 Winner, Natus Vincere vs Fnatic Game 2 Winner, and Natus Vincere vs Fnatic in the best of three. Three contracts, three books, three different outcomes, and string similarity between any two of them is close to total. A matcher gated at 95 percent will happily pair Game 1 against Game 2 and report a spread that is not a spread, it is the difference between two genuinely different questions.
Run the inverse case and the dial fails the other way. The markets tab carries a row reading whether the Fed decreases interest rates by 25 bps after the September 2026 meeting. A second venue might list the same economic event as a rate cut in September with a different construction entirely. Same event, same resolution in practice, similarity score somewhere down in the sixties. That is a real match your gate just discarded.
So the errors you fear most sit above your threshold and the opportunities you want sit below it. Once you have seen that, the checklist writes itself. Every proposed pair needs a human to read both sets of criteria to the end and confirm agreement on the resolution source, the cutoff timestamp including its timezone, whether the question resolves on a date or by a date, the treatment of a tie or a void or a postponement, the scope of who or what qualifies, any early resolution clause, and the unit of measurement where one exists. Any single disagreement in that list turns the pair from an arbitrage into a directional position with a hedge that will not pay.
Measuring the false positive rate rather than arguing about it
The number you need is not available by inspection, so measure it once and then maintain it. The protocol is unglamorous and takes about a day.
- Pull a sample of proposed pairs at each threshold you are considering, sweeping something like 60, 70, 80, 90 and 95 percent. A hundred pairs per bucket is enough to see the shape.
- Define the label before anyone labels anything. A pair is a true match only if a position in both legs would be flat with respect to the underlying event under every resolution path both venues describe. Anything short of that is a false positive, including pairs that are nearly right.
- Have two people label independently and adjudicate the disagreements, because the disagreements are the interesting part and they will tell you which criteria field is doing the damage.
- Sample the rejected pairs at each threshold as well. Precision without recall loss is half a picture, and the discarded true matches are where the strategy's actual capacity went.
- Record the false positive rate by category. Sports and esports series will not behave like policy or macro questions, and a single blended number will hide that.
What you will almost certainly find is a curve that improves as you raise the gate and then flattens or reverses at the top, because the near-duplicate template failures live up there. Whatever shape you get, you now have a measured input rather than a default, and you have the sample to show a risk committee.
The threshold is a queue sizer, the review is the control
Set the dial to whatever produces a review queue your analyst can actually clear, and stop asking it to do risk management. At 80 percent against 1,000 displayed rows the count was 79. If one person can read 79 pairs of resolution criteria in a session, that setting is workable. If they cannot, either raise the gate or add a person, but do not let the queue exceed the review capacity, because a queue nobody clears becomes a queue somebody skims.
Then invert the usual scrutiny. Pairs scoring above 95 percent go into a quarantine that gets more attention, not less, because that band is where the template collisions concentrate. A pair whose two question strings are nearly identical should have to prove it comes from two different venues and two different rule sets before anyone sizes it.
The desk trades from an approved list, never from the live feed. A pair enters that list only after review, and it leaves automatically if either venue amends the question text or the rules, which means someone has to watch for amendments. That watch is the part most implementations skip and it is the part that generates the worst incident, because the pair was genuinely matched on the day it was approved.
The match record and the conversation you will have later
Each approved pair carries a record, and the record exists for a specific future conversation. Both question strings in full, both venue identifiers, both resolution sources, both cutoff timestamps with timezones, both settlement dates, the reviewer, the adjudicator where there was a disagreement, the date, and the similarity score at approval. One page, filed, immutable.
When a pair resolves badly, and one will, that record is the difference between an explanation and an apology. It lets you say precisely which criteria field diverged, that the divergence was not visible at review, and what has changed in the checklist since. Without it the only honest answer available to you is that the model matched them and nobody checked, which is the answer that ends the mandate rather than the position.
Report the control monthly alongside the P&L. Pairs proposed, pairs approved, approval rate, median time to review, pairs rescinded after approval and why, and the P&L attributable to rescinded pairs. That last line is the one that tells you whether the threshold is set correctly, and it is the only line in the whole exercise that a threshold setting on its own could never have produced.