Abstract
Crash data analysis is key to improving road safety, but imbalanced data challenges accurate predictions for severe crashes and often leads to biased outcomes. This study investigates crash severity among young drivers aged 17–24 in England using data collected between April 2019 and February 2022. A standard classification and regression tree model is compared with random undersampling of the majority class CART. Although the undersampled model yields slightly lower overall accuracy, it performs better at identifying severe crashes. Vehicle type and vulnerability, numbers of vehicles and casualties, urban or rural setting, vehicle manoeuvres, dynamic factors, and time-related influences significantly affect young-driver injury severity.
Interactive research sketch
Explore the mechanism
Conceptual interaction based on the study theme—not a reproduction of the reported statistical model.
Illustrative only. Consult the paper for methods, assumptions, uncertainty, and validated results.
Access and rights
Copyright-aware discovery
This record links to an openly available paper or preprint through its DOI. Reuse remains governed by the licence stated at the destination.