Timing and Speed-Measurement Systems for Linear Sprint Testing in Sport: A Systematic Review of Construct Non-Equivalence, Measurement Properties, and Reporting Standards
Peng C, Ma X, Mi J. · Research Square · 2026
Abstract
Background: Linear sprint testing is central to athlete monitoring, talent identification, return-to-play decisions, and research. In high-performance settings, however, practically meaningful changes in sprint time or maximal velocity can be small. Measurement error from timing and speed-measurement systems, start procedures, triggering events, sampling rates, and data processing can therefore be large enough to alter interpretation.
Objectives: This systematic review examined construct non-equivalence in sprint-test outcomes across timing and speed-measurement systems and protocols; mapped validity, agreement, reliability, sensitivity, and methodological error; and developed minimum reporting standards for future sprint-testing studies.
Methods: Records were searched in PubMed, Web of Science, SPORTDiscus via EBSCOhost, and IEEE. Eligible studies evaluated human linear sprint or running-speed tests using manual timing, fully automatic/photo-finish timing, optical timing gates, start-trigger systems, video or smartphone applications, radar or laser systems, GPS/GNSS/local positioning systems, inertial/wearable sensors, motorized sprint-resistance systems, or related approaches. Data were extracted at study, system-comparison, metric, and reporting-quality levels. Outcomes included accuracy, agreement, relative reliability, absolute reliability, sensitivity, methodological error, protocol features, and reporting completeness. Owing to heterogeneity in devices, reference standards, sprint tasks, triggering events, and statistical reporting, the synthesis was primarily narrative and evidence-map based.
Results: The search identified 1258 records. After duplicate removal, 648 records were screened. One hundred and fifty full-text reports were assessed, of which 131 were included in the evidence synthesis. Of these, 90 were core measurement-property studies and 41 were peripheral or methodological studies. The 131 included reports yielded 197 system-level comparisons. Evidence was concentrated in GPS/GNSS/local positioning systems (61 comparisons), timing gates/photocells (47), radar/laser systems (28), inertial or wearable systems (18), video/app/computer-vision
methods (16), manual timing (10), other or multiple systems (10), and motorized or ergometer-based systems (7). Linear sprint splits and acceleration tasks dominated the literature (120 comparisons). However, similarly named outcomes were frequently generated from different start definitions, trigger events, body references, device geometries, sampling rates, filtering choices, and peak-speed definitions. Among applicable reports, reporting was strong for trial and score-selection rules, start distance or posture, device setup, device model, reference standard, and raw or summary
methods (16), manual timing (10), other or multiple systems (10), and motorized or ergometer-based systems (7). Linear sprint splits and acceleration tasks dominated the literature (120 comparisons). However, similarly named outcomes were frequently generated from different start definitions, trigger events, body references, device geometries, sampling rates, filtering choices, and peak-speed definitions. Among applicable reports, reporting was strong for trial and score-selection rules, start distance or posture, device setup, device model, reference standard, and raw or summary results, but weaker for details needed to judge decision usefulness: limits of agreement (18%), smallest worthwhile change or minimum detectable change (22%), bias or mean difference (34%), 95% confidence intervals (42%), and single- versus dual-beam specification (42%). Metric-level evidence was most frequently assigned to agreement and absolute reliability domains, and reliability distributions showed generally high ICC values but more dispersed CV values, reinforcing the need to interpret relative reliability alongside absolute error and agreement.
Conclusions: Sprint-test outcomes are often construct-non-equivalent across systems and protocols; therefore, validity, agreement, reliability, sensitivity, and reporting standards must be interpreted together before sprint data can guide training or research decisions. No timing or speed-measurement system should be considered universally interchangeable. Manual timing, timing gates, radar/laser, GPS/GNSS, video, smartphone, inertial, and motorized systems can each be useful, but only when the measured event, body reference, acquisition settings, processing pipeline, and error statistics match the intended decision. Future studies should adopt minimum reporting standards that make sprint outcomes repeatable, interpretable, and not mistakenly compared across non-equivalent constructs.