The Injury That Data Missed: Why More Monitoring Has Not Made Injuries More Predictable
Prefer listening? Play the audio version:
This is part of our series on the technologies actually changing elite sport in 2026. Earlier deep-dives examined AI injury prediction, wearables, computer vision and recovery technology. This piece steps back to ask a harder question: if elite sport has never collected more data on its athletes, why does the breakdown nobody saw coming remain such a familiar experience inside a performance department? As ever, this is not a manual written from the touchline. It is a map of where the evidence is strong, where it is thin, and what a performance department should ask before trusting a system that did not warn it.
Every performance department has a version of the same story. An athlete is monitored daily. Their GPS output sits within range, their wellness scores look unremarkable, their readiness number is green. Then a hamstring goes in the warm-up, or a soft-tissue injury ends a season, and the post-mortem produces the same uncomfortable finding: nothing in the data flagged it. The instinctive conclusion is that the monitoring was inadequate, that with one more sensor or one more model the signal would have been caught. The evidence suggests a more difficult truth. The problem is frequently not that the data was insufficient. It is that injury, as a phenomenon, does not behave the way a monitoring dashboard implicitly assumes it does.
This article makes one argument. The expansion of athlete monitoring over the past decade has been real and, in important respects, valuable. But the belief that more data should mean more predictable injuries rests on a model of causation that the injury-science literature has spent twenty years dismantling. Understanding why the data misses injuries is not an argument against collecting it. It is an argument for holding it to its actual capabilities rather than its marketed ones.
The prediction problem is worse than most departments assume
Start with the most rigorous assessment of where injury prediction actually stands, because it frames everything else. In 2022, Bullock and colleagues conducted a systematic review of every published musculoskeletal injury-prediction model in sport they could identify, evaluating them against PROBAST, the standard instrument for assessing risk of bias in prediction-model studies (Bullock et al., 2022; Wolff et al., 2019). The review covered thirty studies and 204 separate models. Its central findings were stark: every study had developed a model, and not one had externally validated it. The overwhelming majority were rated at high or unclear risk of bias, and the authors’ conclusion was that no model could be recommended for use in practice.
External validation is not a technicality. A model that has only been tested on the data it was built from has not been shown to work anywhere other than the specific squad, season and conditions it was trained on. Until it is tested against a genuinely independent population, its real-world performance is unknown. The Bullock review found that the entire published literature, at the time, sat below that threshold.
There is also a base-rate problem sitting underneath the modelling problem, and it is the reason this is hard in principle and not only in practice. Serious soft-tissue injuries are rare events at the level of an individual athlete in an individual week. Rare events are difficult to predict usefully even when a model appears to discriminate well in a dataset, because when the underlying rate is low, a model can look impressive in aggregate and still produce far more false alarms than true catches. Flag every athlete with a slightly elevated score and you will be right occasionally and wrong constantly, and a staff that is wrong constantly stops being listened to. The question for a performance department is therefore not only whether a model has a respectable area under the curve, but whether its false positives and false negatives are tolerable for the specific decision it is being asked to support. Resting a player is a different cost from ordering an extra screening, and a model useful for one may be useless for the other.
This is the context a performance director should hold when a vendor presents a prediction tool. The issue is not that injury prediction is impossible in principle. It is that the published evidence base, assessed by its own field’s standards, has not yet demonstrated a model that works reliably outside its development setting. When a dashboard fails to flag an injury, it is behaving consistently with what the methodological literature would predict, not malfunctioning.
Why injury resists prediction: the causation problem
The deeper reason injuries evade monitoring is conceptual, and it predates the current technology entirely. In 2007, Meeuwisse and colleagues published a dynamic, recursive model of injury etiology that has since become one of the most influential frameworks in the field (Meeuwisse et al., 2007). Its central insight is that injury is not the linear product of a single risk factor crossing a threshold. It is the outcome of a complex interaction between intrinsic factors, extrinsic factors and an inciting event, where each repeated exposure to sport alters the athlete’s risk profile, with or without an injury occurring.
The word that matters is recursive. In the Meeuwisse model, an athlete who trains and is not injured is not simply unchanged; the exposure has adapted them, sometimes building resilience, sometimes accumulating asymptomatic microtrauma that lowers their threshold for the next exposure. Risk is therefore not a static quantity a sensor can read on a given morning. It is a moving target shaped by a history of interactions, many of which leave no measurable trace until the moment they matter.
This is the conceptual problem underneath every monitoring system. Most dashboards present risk as a number that can be measured today and acted on today. The etiology literature describes risk as an emergent property of a dynamic system, which is a fundamentally different kind of thing to predict. A single green readiness score on the morning of an injury is not necessarily a measurement failure. It can be an accurate reading of one moment in a system whose relevant state was assembled elsewhere.
The metrics that look like risk but are not
The recursive nature of injury would be merely a theoretical caution if the metrics most departments rely on were robust proxies for it. Several of them are not, and this is where the gap between what a number appears to measure and what it actually measures becomes consequential.
Consider the acute:chronic workload ratio, which for years sat at or near the centre of load-based injury monitoring and still survives in the logic of many load-management conversations, and in some cases in the feature set of risk tools. Impellizzeri and colleagues documented that the ratio suffers from mathematical coupling, that its statistical construction can generate spurious correlations, and that no study had properly attempted to estimate the causal effect the metric is used to imply (Impellizzeri et al., 2020). Their conclusion was direct: there is no evidence supporting its use for training recommendations aimed at reducing injury risk. A follow-up paper argued the underlying theory should be dismissed altogether (Impellizzeri et al., 2021).
The point is not that training load is irrelevant to injury; it plainly matters. The point is that a single ratio was widely treated as if it captured the causal structure of injury risk, and the evidence did not support that. Dropping a more sophisticated algorithm on top of a contested input does not resolve this; it relocates the flaw a layer deeper, where it is harder to inspect.
The translation gap: where good data fails quietly
Even where the underlying measurements are sound, monitoring frequently fails for a reason that has nothing to do with sensors or models. It fails because the data is never converted into a decision.
This is the least technical failure in the category and, in our reading of how monitoring programmes actually live or die inside clubs, among the most common. Monitoring tends to work where a specific person’s job is to translate device output into coaching, medical and load decisions, and it tends to decay where the technology is treated as self-explanatory. The most accurate sensor in the world produces an expensive dashboard nobody opens if no one in the building owns the interpretation.
The failure mode this produces is familiar. A readiness score swings for reasons no one can explain on a given morning, staff lose confidence in it, and within a season the monitoring decays into a compliance exercise that no longer informs decisions. The data is still being collected. It has simply stopped being read in a way that changes what anyone does. When an injury then occurs, the post-mortem sometimes finds that the relevant signal was present in the data, but unattended, because the system that was supposed to surface it had quietly stopped functioning as a decision tool months earlier.
This gives a department a sharper test than data quality alone. A monitoring programme should be judged by decision traceability: which alert changed which conversation, which assessment, which session, which loading decision, and who owned that action. If a system cannot answer that question for the past month, it is not a decision tool. It is an archive.
There is a measurement-quality version of this problem too. Subjective wellness questionnaires, a backbone of daily monitoring, depend entirely on honest, consistent self-report. Where compliance is partial, or where athletes learn which answers keep them in the side, the dataset degrades in ways that are invisible on the dashboard but fatal to any analysis built on it. The number still appears. What it represents has changed.
The move the evidence actually supports: from prediction to prevention
If the argument stopped here it would be a counsel of despair, and it should not, because there is a well-evidenced alternative that the prediction conversation has tended to crowd out. The failure of individual prediction does not imply the failure of injury prevention. It implies that the two have been confused.
Prediction asks which athlete will be injured, and the evidence says no validated model answers it. Prevention asks what reduces injury rates across a squad, and there the evidence is considerably stronger. Injury-prevention programmes built around the Nordic hamstring exercise, for instance, have been shown in meta-analysis to roughly halve hamstring injury rates in footballers, with a pooled injury rate ratio of 0.49 (Al Attar et al., 2017). That is a substantial, replicated effect achieved without knowing in advance which player was going to tear. It works at the level of the population, not the individual forecast.
This is the practical inversion the evidence supports, and it is uncomfortable because it is less glamorous than the dashboard. Stop trying to identify who will break, and apply to everyone the interventions that lower the rate at which anyone does. A department that spends its budget on a predictive model with no external validation, while its squad-wide eccentric strength programme runs at partial compliance, has inverted its own evidence base. The unglamorous intervention has the numbers. The glamorous one has the interface.
What the data can and cannot do
The honest synthesis is that athlete monitoring is genuinely useful for the things it is actually good at, and routinely oversold for the thing it is least able to do. It is well suited to describing what has happened: how much load an athlete has accumulated, how their sleep and resting physiology are trending, how their output compares to their own baseline. These are real capabilities, and used carefully they support better decisions about scheduling, recovery and individual management.
What the evidence does not support is the next step that marketing routinely implies: that the same data can reliably forecast, at the level of the individual athlete in the individual week, who is about to get hurt. The systematic-review evidence says no validated model yet does this. The etiology literature explains why it is so hard. And the metric-level critiques show that some of the inputs feeding these forecasts are weaker than their users assume.
This reframes the injury the data missed. It was not necessarily missed because the monitoring was inadequate. It was missed because prediction at that resolution is a capability the evidence does not show any current system reliably has, and because the conditions for injury were assembled, in the recursive sense, out of interactions that no single reading was positioned to detect.
The position a department can hold
None of this is an argument for collecting less data, and it is certainly not an argument that monitoring is worthless. It is an argument for matching expectations to evidence.
Treat monitoring data as a description of state and a prompt to look more closely, not as a verdict that forecasts the future. When a number flags, the appropriate response is to send a person to assess the athlete, not to act on the number as though it were a diagnosis. When a model claims to predict injury, ask what population it was externally validated in; if the honest answer is none, treat its output as exploratory. Ask which decisions your monitoring system actually changes, and who in the building is accountable for making them, because a programme that no one translates into decisions has already failed regardless of how good its sensors are. Put the prevention work that has evidence behind it ahead of the prediction work that does not. And weigh the cost of the next monitoring tool against the cost of the staff member who could interpret the data you already hold, because the evidence repeatedly points to interpretation, not collection, as the binding constraint.
The injury that data missed is not, in most cases, a story about a sensor that was not sensitive enough. It is a story about a field that has been sold prediction and delivered description, and about the gap between those two things that opens up most painfully at the exact moment a healthy-looking athlete breaks down. Understanding that gap will not prevent the next such injury. But it will protect a department from the avoidable error, which is not missing an injury. It is trusting a system to predict injuries at a resolution the evidence has never shown it can reach.
How this series is made, and how to read it: this is editorial analysis, not a practitioner’s memoir and not a systematic review. PERFORM’s pieces are researched and drafted with the assistance of AI tools, then reviewed, edited and fact-checked by our editorial team against primary sources, peer-reviewed literature, clearly labelled preprints, industry reports, league and company announcements, and practitioners’ own published work. Where the evidence is strong we say so; where it is limited we treat it as limited; where a claim comes from a vendor or corporate announcement we treat it as a hypothesis, not proof. The views here are our editorial position, drawn from the published record rather than first-hand experience inside an elite performance department. Where practitioners are named or quoted, those words are their own. Where we couldn’t verify a claim, we left it out. And where you have the hands-on experience we’re writing about, we’d rather hear from you than pretend to it.
References
Al Attar, W. S. A., Soomro, N., Sinclair, P. J., Pappas, E., & Sanders, R. H. (2017). Effect of injury prevention programs that include the Nordic hamstring exercise on hamstring injury rates in soccer players: A systematic review and meta-analysis. Sports Medicine, 47(5), 907–916. https://doi.org/10.1007/s40279-016-0638-2
Bullock, G. S., Mylott, J., Hughes, T., Nicholson, K. F., Riley, R. D., & Collins, G. S. (2022). Just how confident can we be in predicting sports injuries? A systematic review of the methodological conduct and performance of existing musculoskeletal injury prediction models in sport. Sports Medicine, 52(10), 2469–2482. https://doi.org/10.1007/s40279-022-01698-9
Impellizzeri, F. M., Tenan, M. S., Kempton, T., Novak, A., & Coutts, A. J. (2020). Acute:chronic workload ratio: Conceptual issues and fundamental pitfalls. International Journal of Sports Physiology and Performance, 15(6), 907–913. https://doi.org/10.1123/ijspp.2019-0864
Impellizzeri, F. M., Woodcock, S., Coutts, A. J., Fanchini, M., McCall, A., & Vigotsky, A. D. (2021). What role do chronic workloads play in the acute to chronic workload ratio? Time to dismiss ACWR and its underlying theory. Sports Medicine, 51(3), 581–592. https://doi.org/10.1007/s40279-020-01378-6
Meeuwisse, W. H., Tyreman, H., Hagel, B., & Emery, C. (2007). A dynamic model of etiology in sport injury: The recursive nature of risk and causation. Clinical Journal of Sport Medicine, 17(3), 215–219. https://doi.org/10.1097/JSM.0b013e3180592a48
Wolff, R. F., Moons, K. G. M., Riley, R. D., Whiting, P. F., Westwood, M., Collins, G. S., Reitsma, J. B., Kleijnen, J., & Mallett, S. (2019). PROBAST: A tool to assess the risk of bias and applicability of prediction model studies. Annals of Internal Medicine, 170(1), 51–58. https://doi.org/10.7326/M18-1376