Are AI betting predictions accurate?
AI prediction sites advertise 78–87% accuracy. Those numbers are arithmetically incompatible with the existence of bookmaker margin. Here is what accuracy really measures, why hit rate is the wrong metric, and what to demand instead.
Most advertised AI accuracy figures are meaningless, and a few are arithmetically impossible. A model that goes 85% on odds of 1.20 loses money; a model that goes 45% on odds of 2.60 makes it. Accuracy claims without odds, sample size and drawdown attached cannot be evaluated at all — and the sites publishing them almost never attach any of the three.
Why can an 85% hit rate still lose money?
Hit rate answers only one question: what share of published signals won. It says nothing about what those signals paid. Because bookmakers price favourites near certainty, a high hit rate is available to anyone willing to bet short prices.
The break-even hit rate is fixed by the odds. At decimal odds d, you need to win 1/d of the time simply to stand still:
break-even hit rate = 1 / decimal oddsAt 1.20 that is 83.3%. At 2.60 it is 38.5%.| Average odds | Break-even hit rate | Result at 85% hit rate | Result at 45% hit rate |
|---|---|---|---|
| 1.20 | 83.3% | +2.0 u per 100 | loss |
| 1.50 | 66.7% | +27.5 u per 100 | loss |
| 1.90 | 52.6% | +61.5 u per 100 | loss |
| 2.60 | 38.5% | not achievable in practice | +17.0 u per 100 |
The table is the whole argument. 85% at 1.20 returns two units per hundred bets — a rounding error that a single bad week erases. Meanwhile 45% at 2.60 returns seventeen. A published accuracy number without the average price attached tells you nothing about profitability, which is why serious operators quote yield or units, not percentages won.
Why are claims of 78–87% accuracy not credible?
Bookmaker margin is the arithmetic that kills these claims. A typical three-way football market is priced so that implied probabilities sum to roughly 105–108% rather than 100%. That surplus is the margin, and it is subtracted from every bettor before skill enters the picture.
For a service to sustain 80%+ accuracy on prices that pay meaningfully more than break-even, it would have to be finding systematic mispricing on the majority of events it touches, permanently, in markets watched by professional syndicates with far more data. If that were happening, the prices would move.
The mundane explanations are more likely, and all of them are visible in the wild: counting only favourites, quoting a cherry-picked window, resetting statistics after bad runs, or counting a "correct prediction" as any outcome the model listed among several.
A worked example of how this is done in public. One major US picks site ranks its experts by "streak" — defined in its own glossary as the best profitable run of consecutive picks, searched backwards across all possible lengths. On the day we checked, the second-hottest expert on the front page had a record of 99 wins and 101 losses. Nothing was falsified. The metric was simply chosen so that a losing record could be displayed as a winning one.
What does accuracy actually mean for a forecast?
In forecasting, a model is accurate when its stated probabilities match observed frequencies. If a model says 70% a hundred times, roughly seventy of those should happen. That property is called calibration, and it is not measured by counting wins.
Brier score — the mean squared difference between the stated probability and the outcome (1 or 0). Lower is better; 0.25 is what you get by saying 50% every time.
Log loss — penalises confident errors far more harshly than Brier does. A model that says 99% and is wrong is punished enormously — which is exactly the behaviour you want to discourage.
Both are proper scoring rules: they are minimised only by reporting your true belief, so a model cannot improve its score by exaggerating confidence. Hit rate has no such property, which is precisely why marketing departments prefer it. See measuring prediction quality for how to compute both.
What should you demand before believing an accuracy claim?
- Sample size and period. A hundred settled signals is a starting point, not a proof. Ask over how many events and how many days.
- Average odds. Without it, hit rate is uninterpretable. It is a single number; there is no reason to withhold it.
- The full settled record, losses included — not a highlight reel, and not a metric engineered to find the best sub-window.
- Maximum drawdown and the longest losing run. These describe what following the service actually felt like. Profit describes what a spreadsheet felt like.
- Whether statistics have ever been reset. One tipster-verification platform publishes a counter of how many times each seller has wiped their history. It is the cheapest anti-fraud primitive in the category, and almost nobody offers it.
- Whether model output is ever overridden by hand. One large site states plainly that its experts may override the simulation when they feel confident. That is permitted — but it means the published record is not the model's record.
How does CONSENSUS answer the same questions?
We publish the sample size, the period, the average price, the running result in units, the maximum drawdown and the longest losing run on the front page, and every settled signal in the log with its original odds. The statistics have never been reset, and the model output is never overridden by hand.
We do not publish a headline accuracy percentage, because on the current sample it would be the same misleading number everyone else advertises. We also do not yet publish calibration curves or closing-line value: we are not computing them at the required quality, and inventing them would be exactly the behaviour this guide criticises. When the sample supports them, they will appear with their methodology attached.
What we will say plainly: a sample of roughly a hundred settled signals across two months cannot establish an edge. It can only fail to show one. That distinction is the difference between evidence and marketing.
Check any of this against our record
Every signal CONSENSUS publishes carries the bookmaker odds fixed before the event starts and the settled result afterwards — including the drawdowns and the losing runs. The running total is on the front page and every entry is in the log.