False genius simulator
A leaderboard selects the records most favored by both skill and luck. When many equally capable forecasters are ranked after a finite sample, some must appear exceptional. A fresh out-of-sample period reveals how much of that apparent excellence persists.
No bankroll or staking system is being compared. Every forecaster makes the same number of one-unit predictions at the same odd. The experiment isolates selection, multiple comparisons, and regression to the mean.
In-sample brilliance versus fresh evidence
Each dot is the same forecaster in both periods. Horizontal position is the result used to select the stars; vertical position is independent evidence that selection did not see. A stable skill signal forms a diagonal relationship. Pure luck forms a cloud.
The selected stars
| Period-1 rank | Period 1 ROI | Period 2 ROI | Period-2 rank | True EV |
|---|
The final column is information available to the simulator, not to a real observer. It is shown so that a strong-looking record can be compared with the ability that actually generated it.
Interpretation: regression to the mean does not prove that skill is absent. It proves that selecting extreme past results also selects favorable noise. A track record becomes stronger evidence when it persists out of sample, over a sufficiently large number of independent observations, under a credible model of the population.
