The evidence scoreboard
What a lactate threshold estimate is worth
Your watch has never seen your blood. It infers the threshold from pace and heart rate — and when somebody finally checked that inference against fingertip lactate samples, the pace estimates were out by up to a quarter.
We earn a commission when you buy through our Amazon links, at no extra cost to you. It never decides a ranking — the rubric that does is published in full, so you can check us rather than take our word for it. How this works.
The short answer
Lactate threshold is the intensity above which blood lactate accumulates faster than your body clears it. Watches estimate it from pace and heart rate. Measured against blood lactate testing in 100 runners, threshold pace estimates carried mean errors of 12.7% to 25.8% and all three devices overestimated.
What the threshold actually is
Lactate is produced continuously, at rest and at every exercise intensity, and it is continuously cleared and reused as fuel. What changes with intensity is the balance between those two rates.
At easy intensities, production and clearance keep pace with each other and blood lactate stays near resting levels. Push harder and there is a first inflection — often called LT1 or the aerobic threshold — where concentration begins a sustained rise above baseline. Push harder still and there is a second, sharper inflection, LT2, above which clearance can no longer keep up at all and concentration climbs until you stop. When a coach or a watch says "lactate threshold" without qualification, LT2 is usually what is meant.
The reason anyone cares is that LT2 is one of the better-established markers of the intensity you can sustain for a long effort, and it moves with training in a way that maximum heart rate does not. It is a genuinely useful number. That is exactly why it matters how it was obtained.
How it is measured properly
The reference method is a graded exercise test. You run on a treadmill at a speed that steps up on a fixed schedule — in the study below, 1.2 km/h every five minutes — and at the end of each stage a technician takes a fingertip blood sample and runs it through a lactate analyzer.
That produces a curve of blood lactate against speed, and the threshold is then identified from the shape of that curve. There are several competing methods for doing so. The one used in the validation below is the modified Dmax method, which fits a curve through the points and locates the position furthest from a straight line drawn between two reference points on it.
Two things are worth noticing. The measurement involves actual blood, and even given the blood, identifying the threshold requires a choice of analysis method that reasonable people disagree about. The reference standard is itself a model — a much better-grounded one than a watch's, but not a physical constant.
What your watch does instead
A watch has no access to blood lactate and never will. What it has is pace, heart rate, and the relationship between them over a guided test protocol — typically a progressive run of twenty minutes to an hour, following on-screen instructions. It looks for the point at which heart rate begins to drift upward disproportionately relative to pace, and it calls that the threshold.
The reasoning is sound. Heart rate and lactate do rise together as intensity climbs, and the deflection point in the heart-rate response does tend to occur near the lactate inflection. It is a proxy with real physiological logic behind it. The question is how well the proxy tracks the thing it stands in for, on an individual runner, on a given day.
The measurement, and it is a recent one
Lu et al. (2025) in Frontiers in Physiologyput 100 recreational runners through a treadmill graded exercise test with fingertip blood lactate sampling, then had them perform each watch's own on-device threshold protocol on a 400 m track. Three devices were tested: a Huawei GT Runner, a Garmin Forerunner 265 and a COROS PACE 3.
First finding: a large share of the tests simply failed
Before accuracy, completion. Participants followed each device's protocol correctly and the device still declined to produce a threshold estimate a substantial fraction of the time.
| Device | Attempts | Successful estimates |
|---|---|---|
| Huawei GT Runner | 100 | 78 (78%) |
| Garmin Forerunner 265 | 23 | 15 (65.2%) |
| COROS PACE 3 | 17 | 8 (47.1%) |
Note the sample sizes: only the Huawei figure rests on a large group. Twenty-three and seventeen are small, and the authors flag the imbalance themselves. But a success rate under 50% on any device is worth knowing before you plan a training block around a number your watch may never give you.
Second finding: threshold heart rate looks fine in aggregate and poor per runner
For heart rate, none of the three devices differed significantly from the laboratory value at group level. That sounds like a pass, and it is the sentence a manufacturer would quote. The individual errors tell a different story.
| Device | Mean absolute error | MAPE | Correlation with lab (r) |
|---|---|---|---|
| COROS PACE 3 | 8.93 bpm | 5.95% | 0.13 |
| Huawei GT Runner | 10.66 bpm | 6.32% | 0.36 |
| Garmin Forerunner 265 | 11.44 bpm | 7.15% | 0.67 |
The COROS row is the instructive one. It has the smallest average error anda correlation of 0.13, with an R² of 0.02 — meaning its estimate explained about two percent of the variation in the runners' actual thresholds. A device can land close on average while carrying almost no information about which individual runner it is looking at. That is what a low correlation alongside a small mean error means, and it is the reason we do not report "no significant difference" as a pass.
Nine to eleven beats per minute is also not a small number in practice. It is wide enough to move a threshold-pace session into a genuinely different training zone.
Third finding: threshold pace was materially wrong, and wrong in one direction
| Device | Mean absolute error | MAPE | Direction |
|---|---|---|---|
| Huawei GT Runner | 1.22 km/h | 12.70% | Overestimated (p = 0.01) |
| COROS PACE 3 | 1.93 km/h | 22.63% | Overestimated (trend, p = 0.08) |
| Garmin Forerunner 265 | 2.17 km/h | 25.78% | Overestimated (p < 0.01) |
All three overestimated. That is the practically important part: if you take a threshold pace from your watch at face value and run your tempo sessions at it, the published evidence says you are more likely to be running them too fast than too slow.
The authors also ran equivalence testing, which asks the stricter question — not "is the difference significant?" but "is the agreement close enough to treat the two as interchangeable?" None of the three devices met the equivalence criteria, for heart rate or for pace.
What a wider reading of the literature adds
A 2025 systematic review pooled 13 studies covering 273 participants across Garmin, Polar and Apple devices, foot-worn IMUs and chest-worn GNSS units. Of the five studies that examined lactate threshold, three found the estimate valid and two did not — which is a fair summary of the state of the field.
The two most quotable results in it point in opposite directions, and both are worth carrying:
- In trained runners, one study found speed at threshold acceptably valid — MAPE 7.52% with a concordance correlation of 0.79 — while the heart rate at threshold from the same device fell below the validity cutoff.
- Another found threshold pace 11.96% lower than a field test, while threshold heart rate came within 1.71% and did not differ significantly.
So across studies the direction of the error is not even stable — one found watches overestimating pace by a wide margin, another found them underestimating it. What is stable is that the estimate is noisy at the individual level and that fitness level appears to matter: the review concluded the estimates are reasonable for recreational athletes and uncertain for elite ones.
The review's own limits are worth stating too, because we hold pooled evidence to the same standard as everything else here. It was screened and extracted by a single reviewer, the individual studies were small, none of them tested current-generation watches, and the undisclosed algorithms make cross-device comparison impossible even in principle.
Why this matters more than it first appears
A lactate threshold estimate is not a standalone curiosity. On most platforms it is an input to your training zones, and through the zones it feeds the training-load model, which feeds the status label you actually read.
An error of nine to eleven beats in the threshold shifts every zone boundary derived from it. Sessions get filed in the wrong bucket, load is weighted wrongly, and the resulting status reads confidently while resting on a calibration nobody checked. This is the same compounding problem as everywhere else in this category: the visible output is several transformations away from the last thing that was actually measured.
How to use the number anyway
Treat it as a starting hypothesis, not a prescription. It puts you in the neighborhood. Then test it: a threshold effort you can genuinely hold for the target duration, at an intensity that feels like it should be sustainable and just barely is, tells you more than the watch does.
Prefer the heart-rate figure to the pace figure. On the largest study available, pace errors were roughly two to four times the size of heart-rate errors in percentage terms, and every device overestimated pace. Heart rate is closer to what the device actually observes.
Fix the input first. The Garmin protocol in the study required an ECG chest strap; the other two ran on wrist optical sensing, which the authors flagged as susceptible to motion artifact during running. If the threshold detection is being fed by a wrist sensor during changing intensity, that is the least reliable input this category has.
Retest before you conclude you have improved. Day-to-day threshold variation is real, and it sits on top of a device error of several percent. One higher number after a training block is not evidence of adaptation.
Never compare across brands. Different protocols, different detection algorithms, different analysis methods, all undisclosed. A COROS threshold and a Garmin threshold are not two estimates of one quantity.
The record
What has actually been measured
| Device | Metric | What was actually measured | Evidence |
|---|---|---|---|
| Garmin Forerunner 265 | Lactate threshold pace | MAE 2.17 km/h, MAPE 25.78%, significant overestimation (p < 0.01), n = 15 successful of 23 [1] | Peer-reviewed validation |
| Garmin Forerunner 265 | Lactate threshold heart rate | MAE 11.44 bpm, MAPE 7.15%, r = 0.67; no significant difference from lab, but failed equivalence testing [1] | Peer-reviewed validation |
| COROS PACE 3 | Lactate threshold pace | MAE 1.93 km/h, MAPE 22.63%, overestimation trend (p = 0.08), n = 8 successful of 17 [1] | Peer-reviewed validation |
| COROS PACE 3 | Lactate threshold heart rate | MAE 8.93 bpm, MAPE 5.95% — but r = 0.13 and R² = 0.02, so the estimate carried almost no individual information [1] | Peer-reviewed validation |
| Huawei GT Runner | Lactate threshold pace | MAE 1.22 km/h, MAPE 12.70%, significant overestimation (p = 0.01), n = 78 successful of 100 [1] | Peer-reviewed validation |
| Garmin fēnix 6 (trained runners) | Speed at lactate threshold | MAPE 7.52%, CCC 0.79 — judged valid; heart rate at threshold from the same device fell below the CCC 0.70 cutoff [2] | Peer-reviewed validation |
| Current-generation watches (fēnix 8, Forerunner 970, PACE 4) | Lactate threshold | No validation study exists — the systematic review notes explicitly that none of the reviewed studies tested current models | No validation study exists |
Questions
Frequently asked
How accurate is a watch lactate threshold test?
On the largest published comparison, threshold pace carried mean absolute errors of 12.7% to 25.8% against blood lactate testing, with every device overestimating. Threshold heart rate was closer at roughly 6-7%, or 9-11 beats per minute, but no device met statistical equivalence with the laboratory result.
Can a watch measure lactate?
No. No consumer wearable measures blood lactate. Every watch figure is inferred from the relationship between pace and heart rate during a guided protocol, looking for the point where heart rate drifts up disproportionately. The reference method requires fingertip blood samples analyzed at each stage of a graded test.
Why did my watch fail to give me a lactate threshold?
It is common. In the 2025 validation, tests failed even when participants completed the protocol correctly — the estimate came back on 78% of Huawei attempts, 65.2% of Garmin attempts and 47.1% of COROS attempts. A failed test is the device declining to guess, which is arguably the honest outcome.
Should I use the pace or the heart rate from the test?
The heart rate. Across the study, pace errors were two to four times larger than heart-rate errors in percentage terms and all three devices overestimated pace, meaning threshold sessions run to the watch pace would be run too hard. Heart rate is nearer to what the device actually observes.
Is a lab lactate test worth paying for?
If threshold-based training is central to how you train, it removes the largest source of error in the chain, and every zone and load figure downstream is calibrated from it. Bear in mind that even the lab result depends on which analysis method is used to locate the threshold on the curve, so it is a much better model rather than an absolute constant.
Show your working
Sources
Every one of these was read in full before it was cited. Where we could not access a paper, we do not quote its numbers.
- [1]Peer-reviewed validation2025Lu, Cui, Zhu, Wu, Xing, Pan & Shen, “Validity of smartwatch-derived estimates of lactate threshold heart rate and pace compared to graded exercise testing”, Frontiers in Physiology 16:1621996
DOI 10.3389/fphys.2025.1621996. 100 healthy recreational runners (61 male, 39 female, 28.0 ± 9.6 years) against a treadmill graded exercise test with fingertip blood lactate (Biosen C-line), threshold determined by the modified Dmax method. Device counts were unequal: Huawei GT Runner n = 100, Garmin Forerunner 265/265s n = 23, COROS PACE 3 n = 17. Authors’ stated limits: unbalanced samples, indoor reference against outdoor watch tests, day-to-day threshold variability, and different heart-rate acquisition between devices — the Garmin protocol required an ECG chest strap while the other two used wrist PPG — plus undisclosed manufacturer algorithms.
- [2]Peer-reviewed validation2025Železnik Mežan, “Accuracy of wearables for determining the maximal oxygen uptake and lactate threshold: a qualitative systematic review”, Frontiers in Sports and Active Living 7:1707991
DOI 10.3389/fspor.2025.1707991. 13 studies, 273 participants (181 men, 92 women), covering Garmin Forerunner 230 and 245, fēnix 3, 6 and 7, Polar V800, Apple Watch Series 7, foot-worn IMUs including Stryd, and chest-worn GNSS. Authors’ stated limits: screened and extracted by a single reviewer, small per-study samples, none of the studies tested the latest smartwatch models, undisclosed algorithms prevent cross-device comparison, and the pooled population skewed to 18-30 year olds.
- [3]Peer-reviewed validation2025“Impact of Anatomical Placement on the Accuracy of Wearable Heart Rate Monitors During Rest and Various Exercise Intensities”, Sensors (Basel)
n = 28 (14 male, 14 female). Polar H10 chest strap as the reference. Compared Polar Verity Sense (forearm), Garmin Forerunner 55 (wrist) and three simultaneous Whoop 4.0 units at wrist, forearm and upper arm, across rest, cycling warm-up, burpees and a modified Bruce treadmill protocol.
Read next
Related
- What VO2 max on a watch is worth
The other laboratory value your watch estimates, with its own measured error.
- What training load actually measures
The model your threshold estimate is quietly calibrating.
- Wrist vs chest strap
The input error underneath every heart-rate-derived estimate.
- What running power actually measures
The other pacing number, and why two devices disagree by 82 watts.
- Most accurate heart rate monitors
Fix the measurement the threshold test is built on.