Track record significance calculator
An exact test of a win-loss record against the break-even rate the price demands, with the interval, the posterior, and the separate verdict the money gives.
Worked example — Exact two-sided p-value against break-even: 0.2722. 58 wins, 42 losses at -110 (58.0% win rate): indistinguishable from noise
Looks like an edge (isn’t): p=0.27 against break-even, no evidence despite a 16-point win margin
Advanced options 3
Know the price on both sides? The de-vigged fair rate is a harder null than break-even; see the three de-vig methods on the vig calculator.
The price needs 52.38% to break even. Out of the same 100, it would have taken 63 wins to clear the 95% level.
The test treats every bet as independent with the same true win probability, and the price as an accurate description of what you were paid. Exact two-sided p-value by the method of small p-values: the total probability of every outcome no more likely than this one, under the null. No normal approximation and no continuity correction.
Wilson inverts the score test rather than substituting the observed rate into the standard error, so it stays inside 0 and 100% at every input. At 0 wins or a perfect record Wald collapses to a zero-width interval and calls it certainty.
Beta(1.00, 1.00) prior updated to Beta(59.00, 43.00). The sceptical option centres the prior on the null with the weight of 50 prior bets, which is roughly the position of a reader who has seen a great many records that looked like this one and were nothing. Note that a 87% chance of having an edge is not the same claim as a proven edge, and is not a licence to size as though the rate were the true one.
Per-bet profit is 0.909 units on a win and -1 on a loss, so the per-bet standard deviation is 0.942 units: variance is set by the price, which is why the same return on turnover is far harder to prove at long prices. The two tests agree here. Switch the profit figure to entered units and put in what you actually won, and they can stop agreeing.
Wins needed out of 100 bets to prove an edge, by price
At a fixed sample of 100 bets, the number of wins it takes to clear the 95% level rises with the price, because a heavier price demands a higher break-even rate to begin with.
| Price | Break-even rate | Wins needed of 100 |
|---|---|---|
| -110 | 52.38% | 63 |
| -120 | 54.55% | 65 |
| +100 | 50.00% | 61 |
| +150 | 40.00% | 51 |
Is a 58-42 record proof of an edge?
No. At -110 the exact two-sided p-value against break-even is 0.2722, meaning a coin-flip bettor with no edge at all would produce a record at least this lopsided about one time in four. A 16-point win margin sounds decisive and is not: the honest reading of the same 100 bets is that the true rate sits somewhere between 48.2% and 67.2%, a range that contains both a strong edge and a losing strategy.
Out of the same 100 bets, 63 wins would have cleared the 95% level. Five bets separate a record that proves nothing from one that proves something, which is a fair measure of how little 100 bets can carry.
The null tested above is the break-even rate the price demands, not a coin flip: a bet at -110 is not profitable at 50%, it is profitable at 52.38%. If you know the price on the other side too, the de-vigged fair probability is a harder and better null still, because it asks whether you beat the market's own opinion rather than the market's opinion plus its margin.
Wilson vs Wald, at the extremes
At 0 wins in 30 bets, the Wald interval nearly every calculator prints collapses to 0% to 0%, asserting certainty from a sample that establishes almost nothing. The Wilson score interval at the same input stays inside the range a probability can occupy: 0.0% to 11.4%. Both are computed from the calculator above; this is the failure mode worth seeing rather than describing in the abstract.
What this does not model
- Independence: every test assumes each bet is a fresh draw at a constant true rate. Correlated positions, a drifting strategy, or an adapting market all make a record look more impressive than it is.
- Selection: if you tested twenty ideas and are testing the one that looked best, the p-value on that one is not the p-value of your process.
- Price variance: the per-bet standard deviation on the profit test is estimated from the average price, so a record with widely varying prices carries more noise than the figure shows.
- A p-value is a statement about the data under a hypothesis, never the probability the hypothesis is true; the posterior above answers that second question, and it depends on the prior you pick.
If your own record comes back indistinguishable from noise, find out how many bets your claimed edge would need on the bet sample size calculator. If it clears the level, check the losing run you should still expect on the way with the losing streak calculator: a proven edge and a survivable one are different claims.
What Pro adds here
Bet-by-bet record import, CLV tracking, and rolling significance over a live ledger are Pro features. The launch list sends one email at launch.