Data Mining the Harness Track and Predicting Outcomes

INTRODUCTION Racing, like many other domains including the stock market, trades on publicly available information (i.e. racing histories). Therefore all participants behaving in a rational manner should have an equal chance of success. However, information markets have inequalities through withheld information, a human tendency to discount certain information or weighting it incorrectly. These informational inequalities lead to arbitrage opportunities that can be unlocked using data mining techniques. LITERATURE REVIEW Predictive algorithms have been adopted successfully in racing sports, such as greyhound and thoroughbred racing. These systems use machine learning techniques to train the system on historical data then make predictions on previously unseen data. Highlights of several studies are presented below. The first is a study of greyhound races using ID3 (a decision-tree algorithm) and Back Propagation Neural Network (BPNN) on 100 races at Tucson Greyhound Park (Chen, Rinde, She, Sutjahjo, Sommer, & Neely, 1994). The authors limited themselves to ten race-related variables over a seven race history upon advice from greyhound domain experts: fastest time, win, place and show percentage, break average, finish average, the average finish time over the last three and seven races respectively, competitive grade of the race and a modifier if the horse is competing is a less competitive race. Their system made binary win/no-win decisions for each greyhound based on their historic race data. If a dog was predicted to win the system would make a $2 wager. The ID3 decision tree was accurate 34% of the time with a $69.20 payout ($0.69 excess return per dollar wagered) while the BPNN was 20% accurate with a $124.80 payout ($1.25 excess return per dollar wagered). This disparity in decreased accuracy and increased payout is justified by arguing that the BPNN was selecting longshot winners. As a result accuracy decreases and higher payouts are gained from the longer odds. When comparing their machine learning techniques to track experts, the experts managed a lower 18% accuracy and a payout loss of $67.60 ($0.68 excess loss per dollar wagered). It was speculated that Chen's machine learning system was taking advantage of information inequalities by successfully predicting longshot wagers more often than chance, however, given the black-box nature of BPNN and the difficulties of interrogating its decision-making process, it is hard to be certain. In a follow-up study that expanded the number of input variables to 18, Johansson and Sonstrod used a BPNN on 100 races at Gulf Greyhound Park and found 24.9% accuracy for Win with a $6.60 payout loss ($0.07 excess loss per dollar wagered) [2]. The improvement in accuracy had a corresponding decrease in payout and implies that either the additional variables or too few training cases (449 as compared to Chen's 1,600) confounded their ability to identify longshots. However, the exotic wagers performed better; Quiniela had 8.8% accuracy and $20.30 payout ($0.20 excess return per dollar wagered), while Exacta had 6.1% accuracy and $114.10 payout ($1.14 excess return per dollar wagered). In a third study that focused on using discrete numeric prediction rather than binary assignment, Schumaker and Johnson used Support Vector Regression (SVR) on the same 10 performance-related variables as Chen et. al. [3]. Their study of 1,953 greyhound races employed selective wagering where only the races with the predicted strongest competitors were wagered upon. Of the 505 races selected, their system produced a $0.95 excess return per dollar wagered for Win. In looking at prior research, we discovered a lack of study of machine learners versus the wisdom of crowds. Several studies offered insight between crowds and experts, but none could be found that explored how well a machine learning platform could perform versus crowd wisdom. …

Paper

Full text

PDF

Data Mining the Harness Track and Predicting Outcomes

Semantic Scholar · Business · 2013

Abstract

INTRODUCTION Racing, like many other domains including the stock market, trades on publicly available information (i.e. racing histories). Therefore all participants behaving in a rational manner should have an equal chance of success. However, information markets have inequalities through withheld information, a human tendency to discount certain information or weighting it incorrectly. These informational inequalities lead to arbitrage opportunities that can be unlocked using data mining techniques. LITERATURE REVIEW Predictive algorithms have been adopted successfully in racing sports, such as greyhound and thoroughbred racing. These systems use machine learning techniques to train the system on historical data then make predictions on previously unseen data. Highlights of several studies are presented below. The first is a study of greyhound races using ID3 (a decision-tree algorithm) and Back Propagation Neural Network (BPNN) on 100 races at Tucson Greyhound Park (Chen, Rinde, She, Sutjahjo, Sommer, & Neely, 1994). The authors limited themselves to ten race-related variables over a seven race history upon advice from greyhound domain experts: fastest time, win, place and show percentage, break average, finish average, the average finish time over the last three and seven races respectively, competitive grade of the race and a modifier if the horse is competing is a less competitive race. Their system made binary win/no-win decisions for each greyhound based on their historic race data. If a dog was predicted to win the system would make a $2 wager. The ID3 decision tree was accurate 34% of the time with a $69.20 payout ($0.69 excess return per dollar wagered) while the BPNN was 20% accurate with a $124.80 payout ($1.25 excess return per dollar wagered). This disparity in decreased accuracy and increased payout is justified by arguing that the BPNN was selecting longshot winners. As a result accuracy decreases and higher payouts are gained from the longer odds. When comparing their machine learning techniques to track experts, the experts managed a lower 18% accuracy and a payout loss of $67.60 ($0.68 excess loss per dollar wagered). It was speculated that Chen's machine learning system was taking advantage of information inequalities by successfully predicting longshot wagers more often than chance, however, given the black-box nature of BPNN and the difficulties of interrogating its decision-making process, it is hard to be certain. In a follow-up study that expanded the number of input variables to 18, Johansson and Sonstrod used a BPNN on 100 races at Gulf Greyhound Park and found 24.9% accuracy for Win with a $6.60 payout loss ($0.07 excess loss per dollar wagered) [2]. The improvement in accuracy had a corresponding decrease in payout and implies that either the additional variables or too few training cases (449 as compared to Chen's 1,600) confounded their ability to identify longshots. However, the exotic wagers performed better; Quiniela had 8.8% accuracy and $20.30 payout ($0.20 excess return per dollar wagered), while Exacta had 6.1% accuracy and $114.10 payout ($1.14 excess return per dollar wagered). In a third study that focused on using discrete numeric prediction rather than binary assignment, Schumaker and Johnson used Support Vector Regression (SVR) on the same 10 performance-related variables as Chen et. al. [3]. Their study of 1,953 greyhound races employed selective wagering where only the races with the predicted strongest competitors were wagered upon. Of the 505 races selected, their system produced a $0.95 excess return per dollar wagered for Win. In looking at prior research, we discovered a lack of study of machine learners versus the wisdom of crowds. Several studies offered insight between crowds and experts, but none could be found that explored how well a machine learning platform could perform versus crowd wisdom. …

Similar papers

© 2026 NYSGPT2525 LLC