A backtest reconstructs how a model would have performed on races that have already happened; a live or forward record logs predictions on future races as they occur. A well-designed backtest can test whether an idea generalizes across time, but it is still historical. A prospective record can test current data and execution, but a short run can be distorted by chance. Neither alone proves a durable betting edge.
What “backtested” and “live” results actually mean
| Aspect | Backtest | Forward or live record |
|---|---|---|
| When the evidence is produced | Past races are replayed using a model and historical data. | Predictions are recorded on future races as they happen. |
| Information available | Must be reconstructed for the model’s decision time; historical databases may contain later information. | Uses the data feed available at prediction time, which should be timestamped and retained. |
| Model and rules | Can be tuned repeatedly, creating a risk of overfitting. | Should be frozen for the evaluation period so results are interpretable. |
| Prices and execution | Often depends on assumed or stored odds and may not reflect obtainable prices. | Can record actual offered and taken prices, rejected or partial bets, commission and slippage. |
| What it can test | Historical patterns, model development and performance on a later holdout period. | Prospective generalization and operational behavior under current conditions. |
| Main risks | Information leakage, data snooping, selection bias and unrealistic price assumptions. | Small samples, variance, changing markets and selective reporting. |
A chronological test on races that have already occurred is still a backtest, not a live result. It is generally a more informative historical test than a random split because it asks how a model fares on later data, but it does not show that bets could actually have been placed at the assumed prices.
How to judge whether a backtest is credible
1. Rebuild the information available at decision time
Set the intended prediction time before evaluating the model. For every feature, ask whether it would genuinely have been available at that moment. Exclude results, payouts, final odds, finishing positions and other post-race information, as well as any same-race information derived from outcomes.
This is the central leakage risk: information recorded in a database after a race can accidentally enter feature engineering, preprocessing, model selection or evaluation. Shuichi Sugiura’s 2026 study describes this problem and constrains predictors to information available after entries were finalized but before outcomes were known. It also excludes some same-day variables when their availability or stability at the intended prediction time is uncertain. Read the study.
#1 Best Overall
2. Keep the data in time order
Use earlier races to fit the model, a later period to compare candidate models or set parameters, and a still later period for a final assessment. Do not keep changing features or rules in response to results from that final period; once those results influence development, the period is no longer an untouched test.
Sugiura’s Japanese flat-racing study used 2015–2022 for training, 2023–2024 for validation and January 5, 2025–May 10, 2026 for independent testing. Its test set contained 63,910 horse-level observations from 4,556 races. Those dates and counts describe that study, not a required split or sample size for other racing codes, jurisdictions or models. A separate 2026 race-level study also used a temporal arrangement and reserved its test period for final evaluation. Read the race-level study.
Rank #2
- Ideal for Gifting
- Ideal for a bookworm
- Compact for travelling
3. Challenge historical profit claims
Testing many combinations of odds bands, race types, filters and model settings makes it more likely that one historical slice will look profitable by chance. Compare the model with a relevant market benchmark and a simple baseline, and check whether results hold across time periods rather than hinging on one unusually successful run. Treat subgroups chosen after viewing outcomes as exploratory, not as independent confirmation.
4. Use prices the strategy could have obtained
Betting returns depend on the price available when the strategy would act. Match odds to that decision time, account for exchange commission where applicable, and document non-runners and other race changes. Keep both the trigger price and the price actually obtained; rejected or partially matched bets and slippage can turn a theoretical edge into a different realized result.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
How to evaluate a forward or live record
Before the test starts, freeze the model version, features, selection rules and staking method for a defined period or sample. Log every qualifying prediction, including losing selections, rather than removing entries after seeing the outcome.
- Timestamp of the prediction and the data used.
- Model probability or rating, plus any expected or fair price.
- Price available when the bet was placed and, if relevant, the closing price.
- Actual stake, result and return, including commission where applicable.
- Execution problems, such as an unavailable price, rejection or partial fill.
- Any change to the data feed, model or operating rules.
Paper testing can record selections exactly as if bets were being made without risking money. It helps assess the prediction process, but it cannot establish whether real prices are accessible or reveal execution slippage. A small-stakes execution check may expose those operational issues; it is not a guarantee that future results will be profitable.
Rank #4
Which metrics answer which question?
Prediction quality and betting performance are different. A model may rank likely winners well while giving poorly calibrated probabilities, or produce useful probabilities without generating profitable bets at available prices.
| Question | Useful measures | What they tell you |
|---|---|---|
| Does the model distinguish winners from non-winners? | ROC AUC; PR-AUC where appropriate | How well scores rank positive outcomes. These are not measures of profit. |
| Are its probabilities reliable? | Brier score; log loss | How close probability estimates are to outcomes, with penalties for inaccurate probabilities. |
| Did the betting rules make money in the stated sample? | Number of bets, total stakes, returns, profit, ROI or yield, average odds, maximum drawdown and longest losing run | The observed financial result and some of its downside and volatility. It does not by itself establish a persistent edge. |
State the comparison baseline, such as market-implied probabilities, a margin-adjusted market benchmark where possible, a favourite baseline or a simpler ratings model. Strike rate alone is not enough: its meaning depends on the odds. Report losses and uncertainty alongside returns. There is no universal number of live bets that proves an edge; the relevant sample depends on odds, outcome variance and the consistency of results.
Best Value
- It can be a gift option
- Comes with secure packaging
- Helpful in various ways
What published horse-racing studies can—and cannot—show
Sugiura’s 2026 peer-reviewed study evaluates Japanese flat racing with a time-ordered historical test, not a prospective betting ledger. On its test set, the matched no-theory model reported a win ROC AUC of 0.7543 (95% CI 0.7475–0.7609), compared with 0.7293 (95% CI 0.7224–0.7362) for the augmented current-full model. For the study’s JRA place-rule-compatible outcome, the corresponding AUCs were 0.7513 (95% CI 0.7469–0.7558) and 0.7164 (95% CI 0.7118–0.7212). These are study-specific measures of discrimination, not ROI estimates or evidence of performance in other markets.
The separate 2026 race-level paper evaluates an upset-risk diagnostic and says it was not integrated into horse-level prediction scores. A result for a race-level instability indicator therefore does not demonstrate that a horse-selection model can make money. See its scope and methods.
A 2026 SSRN preprint on French trotting at Vincennes describes chronological evaluation and a retrospective backtest settled at official PMU dividends. A preprint’s simulated historical returns remain distinct from live realized returns and should not be generalized into a claim that racing markets are beatable. Read the preprint.
Practical testing guidance also emphasizes comparing against a market or simpler model and disclosing stability, sample size and execution assumptions. See the guide to testing a horse-racing betting model.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




