Any public leaderboard produces the same statistical artefact. You are shown the maximum of a set of noisy estimates and invited to read it as the best member of the set. Those are different claims. With enough entrants, the top of a board is populated by whichever strategies drew the most favourable noise, and the gap between the displayed number and the defensible number widens as the board fills up.
The Momentum Trading Engine's leaderboard tab ranks deployments and offers Overall Score, Total Return, Risk-Adjusted Return, Win Rate, Alpha and Profit Factor as ordering choices, scoped across all time, crypto or stocks. Risk-Adjusted Return is the one that invites the haircut, because a Sharpe estimate has a known sampling distribution and a maximum over many draws from it has a known expectation.
Two inputs, and both are available
The correction needs the standard error of a single Sharpe estimate and the expected maximum of that many independent estimates under a null of no skill. Multiply them and you have the amount by which the rank-one entry should be discounted before you treat it as evidence of anything.
The standard error of a Sharpe ratio over a sample of length T is approximately the square root of one plus half the squared Sharpe, all divided by T. At capture the engine reported an average Sharpe of 1.58 across nine live strategies and 350 trades backtested in total, which is roughly 39 observations per strategy. Substituting, one plus half of 1.58 squared is 2.248, divided by 39 gives 0.0576, and the square root is 0.240.
That figure alone should slow you down. A Sharpe estimate on this sample length carries a standard error of about a quarter of a unit, so a displayed 1.58 and a displayed 1.10 are inside one standard error of each other. The board is ranking on differences it cannot measure.

The haircut, worked
The expected maximum of N independent standard normal draws grows with the square root of twice the natural log of N. It is a slow function, which is the point. Going from ten entrants to a hundred does not multiply the expected maximum by ten, it adds about forty percent to it.
Take the engine's nine live strategies as the smallest plausible N. Twice the log of nine is 4.39, square root 2.10. Multiply by the standard error of 0.240 and the haircut is 0.50. A rank-one Sharpe of 1.58 deflates to about 1.08.
Now assume a board that has grown to 25 deployments. Twice the log of 25 is 6.44, square root 2.54, haircut 0.61, and 1.58 becomes 0.97. At 100 deployments the square root term is 3.04, the haircut is 0.73, and the surviving Sharpe is 0.85.
That is the finding, and it is not dramatic in either direction. The rank-one entry on a board of a hundred, showing a headline Sharpe of 1.58 on 39 trades, is defensible as a strategy with a Sharpe somewhere just under one. That is a real but ordinary result, and it is nowhere near what the displayed ranking implies. Note that the square root of twice the log is the leading asymptotic term and runs somewhat high at small N, so at nine or twenty-five entrants treat these haircuts as the conservative end.
Which way the independence assumption fails
The correction assumes N independent trials, and that assumption breaks in both directions at once. You need to think about both or you will apply the haircut with false confidence.
It breaks downward because entrants are correlated. If the board's strategies are drawn from eleven entry systems grouped into three families and running on five synchronised timeframes, many entrants are variations on the same idea and the effective number of independent trials is smaller than the headcount. Correlated entrants produce a lower expected maximum, so the true haircut is smaller than the naive N implies.
It breaks upward, and by more, because the visible board is not the trial count. The engine's signal feed labels configurations by generation and number, with entries at capture reading as version two configurations numbered into the twenties. A version two implies a version one. Configuration numbers in the twenties against nine live strategies implies configurations that were tested and did not graduate. The N that belongs in the formula is every configuration ever evaluated, not the subset that survived to appear on a board, and that number is not disclosed.
On balance the second effect dominates, because the discarded population is usually much larger than the correlation discount. When you write this up, state the visible N, state that the true N is larger and undisclosed, and present the haircut as a floor.
The disclosure gap that undermines the whole calculation
There is a problem sitting underneath all of this arithmetic, and it is worth raising before the numbers rather than after. The engine displays an average Sharpe of 1.58 and does not state the period it is computed over. A Sharpe per trade, a Sharpe on daily returns and an annualised Sharpe differ by large multiplicative factors, and the deflation math needs to know which one you have.
It also needs to know what T is. I used 39 because 350 trades across nine strategies is the disclosed arithmetic, but trades are not the natural unit for a Sharpe ratio. If the underlying series is daily returns over three years, T is closer to 750 and the standard error drops to about 0.055, which shrinks every haircut above by more than a factor of four and leaves the rank-one entry looking considerably better.
So the sequence for a real diligence file is to establish the return frequency and window first, then compute the standard error, then apply the haircut. Getting that in writing from the vendor is the highest value question on the list, because the answer moves the conclusion from ordinary to genuinely interesting or the other way, and no amount of care in the later steps compensates for guessing at it.
What to shortlist from instead
Given all of the above, ranking on Risk-Adjusted Return and taking the top entry is the one thing not to do. Two alternatives are better and neither is difficult.
The first is to shortlist the strategies that appear in the top quartile of several orderings at once. The board offers six. A strategy in the top quartile by Risk-Adjusted Return, Alpha and Profit Factor simultaneously has survived three partly independent screens, which is a weaker form of the same idea as requiring statistical significance but is robust to the sample being short.
The second is to rank on the lower bound rather than the point estimate. Subtract one standard error from every displayed Sharpe before you sort. On a 39-trade sample that subtracts 0.24 from everything, which changes no ordering by itself, but it makes the sort visibly a ranking of pessimistic estimates and that framing tends to survive contact with an investment committee far better than a leaderboard screenshot does. At capture the board carried no ranked strategies at all, which means the first person to deploy will occupy rank one on a board of one, and a rank-one finish out of one entrant deserves exactly the haircut that implies.