The evidence scoreboard
What training load actually measures
Every brand has one, none of them means the same thing, and not one has been independently validated. Here is what it is built from.
We earn a commission when you buy through our Amazon links, at no extra cost to you. It never decides a ranking — the rubric that does is published in full, so you can check us rather than take our word for it. How this works.
The short answer
Training load is a single number summarizing how much stress a session imposed, computed from time spent at each heart rate relative to your thresholds. It is arithmetic on heart-rate data, not a measurement — and no independent validation of any manufacturer's load model exists that we could find.
The chain, one link at a time
It helps to see how many steps sit between your heartbeat and the number in the app.
- A sensor detects your heart rate. If it is a wrist optical sensor, the published evidence says this step is reliable during steady work and much less reliable during rapid intensity change.
- The watch compares that against thresholds — maximum heart rate, lactate threshold, or zones — which are themselves usually estimated rather than tested.
- Time in each zone is weighted, typically exponentially so that hard minutes count for far more than easy ones.
- The weighted total becomes a session score, which accumulates into rolling short-term and long-term averages.
- The ratio between those two averages produces the "productive", "maintaining" or "overreaching" label you actually read.
Every link multiplies the uncertainty of the one before it. A load score is a fifth-order derivative of a measurement that, on a wrist, has documented failure modes at exactly the intensities that dominate the calculation.
What the acute-to-chronic ratio is trying to say
The idea underneath most load models is straightforward and sensible: compare what you have done recently against what you are used to doing. Train far above your habitual load and you accumulate fatigue faster than you adapt; train far below it and you lose fitness.
The labels brands attach to that ratio are interpretations of a real concept. Where they differ — and they differ a lot — is in the weighting, the time windows and the thresholds, none of which is published in enough detail to reproduce.
Why the numbers are not comparable across brands
Garmin, COROS, Polar, Suunto and Apple all compute something they call training load or training status, from proprietary models, on different scales. A COROS number and a Garmin number are not two measurements of the same quantity — they are two different opinions expressed in different units.
This matters practically when you switch watches. Your load history does not transfer in any meaningful sense, and a lower number on the new device is not evidence that you trained less.
The evidence position, stated plainly
We could find no independent validation of any consumer training-load model. Not peer-reviewed, not a measured enthusiast test. Nothing.
The closest relevant finding is from the 2025 nocturnal HRV validation, whose authors noted there is "little transparency into what metrics affect each device's own Readiness or Recovery Score" and that these algorithms are periodically updated without notice. That was written about recovery scores, and it applies with equal force here.
So the honest evidence tier for every training-load score on the market is no validation study exists. That is not the same as saying the numbers are useless — it is saying nobody has checked, and any page that tells you one brand's load model is more accurate than another's is making it up.
How to use it anyway
As a relative signal on one device. Consistent arithmetic applied to consistent inputs will still show you when this week was much harder than last. That is genuinely useful and it does not require the model to be validated.
Fix the input before you trust the output. If your heart-rate data comes from a wrist sensor during interval sessions, the load score is built on the measurement with the largest documented error in this category. A chest strap improves everything downstream of it, which is the single highest-leverage change available — the upgrade path is here.
Get your thresholds right. A load model calibrated against an estimated maximum heart rate that is 15 beats wrong will misweight every session you do. Most watches let you set these manually; a field test is more useful than a formula.
Do not let it override how you feel. A number computed from an unvalidated model on top of an inferred measurement should not outrank the fact that you are exhausted.
Questions
Frequently asked
What does training load actually measure?
Time spent at each heart rate during a session, weighted so that harder intensities count for disproportionately more, then accumulated into short-term and long-term averages. It is arithmetic performed on heart-rate data rather than a physiological measurement.
Is Garmin or COROS training load more accurate?
Nobody can answer that, because no independent validation of either model exists that we could find. They use different weightings, windows and thresholds and report on different scales, so the numbers are not two measurements of the same quantity. Anyone ranking them for accuracy is guessing.
Why did my training load drop when I changed watches?
Because it is a different model on a different scale, usually with different threshold settings. Your load history does not transfer between brands in any meaningful way, and a lower number on the new watch is not evidence that you trained less.
Does a chest strap improve training load accuracy?
It improves the input, which is the only part of the chain you can control. Load is computed from heart rate, and wrist optical sensing has its largest documented errors during exactly the rapid intensity changes that dominate a load calculation. Better input does not validate the model — it just stops the model working from bad data.
Should I train by training load?
Use it as one input among several, on a single device, as a relative signal over weeks. Treat the label — productive, overreaching, unproductive — as a suggestion from an unaudited model rather than a finding, and do not let it override how you actually feel.
Show your working
Sources
Every one of these was read in full before it was cited. Where we could not access a paper, we do not quote its numbers.
- [1]Peer-reviewed validation2025Dial, Hollander, Vatne, Emerson, Edwards & Hagen, “Validation of nocturnal resting heart rate and heart rate variability in consumer wearables”, Physiological Reports 13(16):e70527
DOI 10.14814/phy2.70527. n = 13 healthy adults (7 male, 6 female, 33.2 ± 8.6 years) across 536 nights, each wearing an Oura Gen 3, Oura Gen 4, Polar Grit X Pro, Garmin fēnix 6 and Whoop 4.0 simultaneously against a Polar H10 single-lead ECG reference sampled at 1000 Hz. Authors’ stated limits: healthy adults only, no atrial-fibrillation population, proprietary and periodically-updated algorithms, and unequal night counts per device. The Garmin was dropped from the resting-heart-rate comparison because the 30-minute window it uses is not timestamped.
- [2]Peer-reviewed validation2025“Impact of Anatomical Placement on the Accuracy of Wearable Heart Rate Monitors During Rest and Various Exercise Intensities”, Sensors (Basel)
n = 28 (14 male, 14 female). Polar H10 chest strap as the reference. Compared Polar Verity Sense (forearm), Garmin Forerunner 55 (wrist) and three simultaneous Whoop 4.0 units at wrist, forearm and upper arm, across rest, cycling warm-up, burpees and a modified Bruce treadmill protocol.
- [3]Manufacturer specification2026Polar — H10 heart rate sensor published specifications
Manufacturer-stated: electrical (ECG) measurement, Bluetooth LE + ANT+, internal memory for one session up to 30 hours, up to 400 h on a CR2025 coin cell, WR30.
Read next
Related
- What VO2 max on a watch is worth
The other derived metric — this one has actually been measured.
- Fix your heart-rate data
The input everything on this page is computed from.
- HRV accuracy by device
The one recovery input that has been validated across five devices.
- What a lactate threshold estimate is worth
The threshold your load model is calibrated against, and how far off it can be.