How to evaluate a tipster or prediction service
Nine checks that separate a verifiable record from a marketing claim: sample size, average odds, drawdown, statistic resets, verification method and who controls the numbers. With real examples of each failure mode.
Evaluate a tipster the way you would evaluate a fund: demand the sample size, the average price, the full settled history including losses, and the worst drawdown. If any of the four is missing, you cannot assess the service — and in this category, missing numbers are usually a choice rather than an oversight.
The nine checks, in order of how quickly they disqualify
- Is there a full settled record, or only highlights? If you cannot see every published signal with its result, nothing else on the page is evidence.
- What is the average price? Without it, a hit rate is uninterpretable — see why 85% can lose money.
- What is the sample size and over what period? A hundred signals across two months is a starting point. Two thousand across three seasons is an argument.
- Is the maximum drawdown published? Almost nobody publishes it. It is the single most informative number for someone deciding how much to risk.
- Can statistics be reset, and is the count public? Resets turn a history into a selection.
- Who verifies, and who pays for the verification? Verification bought by the seller is marketing with a badge.
- Is the metric defined? Watch for invented metrics like "current streak" that search backwards for the most flattering window.
- Is model output ever overridden manually? If yes, the record belongs to the humans, not the model.
- Does the arithmetic reconcile? Multiply it out. Wins divided by settled signals should equal the stated hit rate.
What does each failure mode look like in the wild?
These are not hypotheticals. Each was observable on a live commercial page while researching this guide.
| Failure mode | What it looks like | What it hides |
|---|---|---|
| Engineered metric | Experts ranked by "best consecutive streak", searched backwards across all lengths | A featured expert with a 99–101 record |
| Arithmetic that does not close | Claimed 78.05% hit rate on 945 tips with 707 wins and 28 draws | 707/945 is 74.8%; no formula reproduces 78.05% |
| Paid verification | A "verified tipster" badge sold to the tipster for €1,000 | That nothing was independently checked |
| Record behind a paywall | A site advertising "full transparency" with accuracy reports at $57/month, public figures dated 2012–2020 | Current performance |
| Profit without drawdown | Yearly profit, strike rate, 30-day and 7-day profit per tipster — no drawdown anywhere | What following it felt like |
| Survivorship in leaderboards | A leaderboard of 2,000+ tipsters ranked by profit | That the top of any large sample looks skilled by chance |
Why does maximum drawdown matter more than profit?
Profit is the endpoint. Drawdown is the path. A service can end a season up ten units having been down thirty in March, and almost nobody who started in February is still following it by then. Followers do not quit on a bad month; they quit on a bad month they were not warned about.
A practical bankroll rule used by the few operators who publish drawdown at all: size your bank at roughly 1.5–2× the historical maximum drawdown of the strategy you are following. If the service will not tell you its worst stretch, you cannot apply the rule, and you are sizing blind. See maximum drawdown.
What does good verification look like?
The strongest verification practice we found in the category comes from tipster marketplaces rather than model operators: results checked in real time against a reference bookmaker feed, a badge that is lost automatically when too many entries cannot be verified, and — crucially — a public counter of how many times the seller has reset their statistics, which active sellers cannot clear.
The second-strongest is mechanical publication gating: a strategy is published only if it clears fixed thresholds, cannot be edited after saving, and is delisted automatically when its drawdown exceeds a stated multiple of the original — regardless of whether it is still profitable. Rules that fire against the operator's own interest are the ones worth trusting.
The same checklist applied to CONSENSUS
| Check | CONSENSUS |
|---|---|
| Full settled record | Yes — every signal, with original odds, in the open log |
| Average price published | Yes, on the front page |
| Sample size and period | Yes, stated next to every figure |
| Maximum drawdown | Yes, in the headline band at the same size as profit |
| Statistics ever reset | Never |
| Verification | Self-published and independently recountable — we do not claim third-party verification, because we have not bought any |
| Metric definitions | Plain: hit rate is wins ÷ settled; result is units at one unit flat |
| Manual override of model output | None |
| Calibration and closing-line value | Not published — we do not yet compute them to a standard worth publishing |
The last two rows are the honest part. We do not have third-party verification and we do not have calibration figures yet. A checklist you can only pass is a sales page, not a checklist.
Check any of this against our record
Every signal CONSENSUS publishes carries the bookmaker odds fixed before the event starts and the settled result afterwards — including the drawdowns and the losing runs. The running total is on the front page and every entry is in the log.