Menu

HomeHow it worksContact UsContact Us

Explainable temporal deep learning for athlete-independent classification of injury-labeled days in competitive distance runners: a methodological benchmark using seven-day training-load histories.

Qiao T, Tian W. · Frontiers in public health · 2026

Abstract (source)

Introduction: This study evaluates whether seven-day training-load histories can support athlete-independent classification of injury-labeled days in competitive distance runners. The

Aim: is methodological benchmarking rather than clinical injury-risk deployment. Research gap Prior sports-injury prediction studies often use repeated athlete-day observations without fully addressing athlete leakage, rare-event imbalance, probability calibration, and the risk that explainability

Methods: may identify recording artifacts rather than physiological mechanisms. Data and

Method: The public dataset contained 42,766 athlete-day observations from 74 runners, including 583 injury-labeled observations (1.36%). Seven external-load variables and three internal-response variables were organized from Day -7 to Day -1. A compact dual-stream multi-scale temporal convolutional network with gated fusion and feature-by-day Integrated Gradients was evaluated using athlete-independent five-fold grouped cross-validation. Logistic regression and XGBoost served as baselines, while three ablations and five prespecified neural-network seeds assessed robustness.

Findings: In the primary analysis, DS-MTCN-IG achieved ROC-AUC 0.596, PR-AUC 0.0204, and Brier score 0.01343. Logistic regression achieved a numerically higher ROC-AUC of 0.619, but the paired athlete-cluster bootstrap difference was not statistically significant (ΔROC-AUC = 0.0227, 95% CI - 0.0180 to 0.0714; p = 0.291). Across five seeds, the ungated dual-stream model achieved the highest mean ROC-AUC among the deep variants (0.604 vs. 0.598 for the gated model), so the gating hypothesis was not supported. In the non-overlapping evaluation, which removed temporal adjacency between successive prediction targets, DS-MTCN-IG achieved a mean ROC-AUC of 0.469 and did not demonstrate above-chance discrimination. This finding indicates that the weak discrimination observed in the full-window analysis may reflect temporal overlap and autocorrelation between adjacent seven-day histories rather than a stable athlete-independent relationship between training-load history and injury-labeled days. Five-seed Integrated Gradients identified contributions from measured internal-response ratings and availability masks, but exact top-input rankings were unstable across seeds.

Conclusion: The study did not demonstrate a robust or practically useful athlete-independent predictive signal. The non-overlapping null

Result: substantially limits interpretation of the full-window estimates and provides no support for clinical, coaching, or automated injury-warning use. The principal contribution is therefore a methodological audit showing how temporal overlap, participant-aware validation, calibration, model comparison, and attribution stability can alter

Conclusions: in longitudinal sports-injury modeling.

Primary studyOpen accessProgramming & Periodization
Read the original →