Predicting Crash Severity on the Hungarian Road Network: An Ensemble Machine Learning Approach With Resampling and Hyperparameter Tuning
Accurately predicting and interpreting crash severity is critical for developing targeted and cost-effective road safety strategies. While machine learning (ML) models have demonstrated considerable promise in this field, two persistent challenges remain: the limited interpretability of ML algorithms and the inherent imbalance of crash datasets, particularly the scarcity of fatal injury crashes. This study addresses these gaps by applying Random Forest, XGBoost, and LightGBM classifiers to traffic crash data from the Hungarian road network, incorporating both traffic and roadway characteristics. Additionally, feature selection strategies, including Pearson correlation, Mutual Information, and RFE with Logistic Regression, were implemented to improve interpretability and reduce redundancy. To counter dataset imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was employed, while model robustness and generalizability were enhanced through systematic hyperparameter tuning using GridSearchCV. Model evaluation, based on accuracy, precision, recall, F1-score, MCC, and per-class G-Mean, revealed that the Random Forest classifier consistently outperformed alternative models, achieving the highest accuracy and discrimination capacity. Feature importance analysis underscored the dominance of traffic volume, horizontal road geometry, and heavy vehicle composition, with the “Annual Average Daily Cross-sectional Traffic” and “Radius of Horizontal Alignment Curve” emerging as the most influential predictors of crash severity. By providing both methodological contributions and context-specific insights, this study offers valuable guidance for transportation authorities in Hungary and beyond, supporting the development of data-driven, precise, and cost-effective interventions aimed at mitigating road traffic injuries and fatalities.
Paper
Full text
Predicting Crash Severity on the Hungarian Road Network: An Ensemble Machine Learning Approach With Resampling and Hyperparameter Tuning
Semantic Scholar · Computer Science · 2025
Abstract
Accurately predicting and interpreting crash severity is critical for developing targeted and cost-effective road safety strategies. While machine learning (ML) models have demonstrated considerable promise in this field, two persistent challenges remain: the limited interpretability of ML algorithms and the inherent imbalance of crash datasets, particularly the scarcity of fatal injury crashes. This study addresses these gaps by applying Random Forest, XGBoost, and LightGBM classifiers to traffic crash data from the Hungarian road network, incorporating both traffic and roadway characteristics. Additionally, feature selection strategies, including Pearson correlation, Mutual Information, and RFE with Logistic Regression, were implemented to improve interpretability and reduce redundancy. To counter dataset imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was employed, while model robustness and generalizability were enhanced through systematic hyperparameter tuning using GridSearchCV. Model evaluation, based on accuracy, precision, recall, F1-score, MCC, and per-class G-Mean, revealed that the Random Forest classifier consistently outperformed alternative models, achieving the highest accuracy and discrimination capacity. Feature importance analysis underscored the dominance of traffic volume, horizontal road geometry, and heavy vehicle composition, with the “Annual Average Daily Cross-sectional Traffic” and “Radius of Horizontal Alignment Curve” emerging as the most influential predictors of crash severity. By providing both methodological contributions and context-specific insights, this study offers valuable guidance for transportation authorities in Hungary and beyond, supporting the development of data-driven, precise, and cost-effective interventions aimed at mitigating road traffic injuries and fatalities.