The evidence scoreboard
What running power actually measures
Cycling power is measured force. Running power is a model — and when a study ran two of the leading models on the same trail at the same time, they disagreed by an average of 82 watts.
We earn a commission when you buy through our Amazon links, at no extra cost to you. It never decides a ranking — the rubric that does is published in full, so you can check us rather than take our word for it. How this works.
The short answer
Running power is a modeled estimate, not a measurement. No sensor measures the force you apply to the ground. On a trail course, a Stryd foot pod and a Garmin watch were each highly repeatable but agreed poorly with one another — a mean difference of 82.2 watts, rising to 109.1 watts uphill.
The difference from cycling power, which is the whole story
A bicycle power meter contains strain gauges in the crank, pedal or hub. When you push, the metal deforms by a tiny amount, the gauges convert that deformation into a voltage, and the device multiplies the resulting torque by angular velocity. Watts come out. It is a direct physical measurement of mechanical work being done on a rigid object, and that is why a cycling watt is a portable unit: 250 W on your bike is 250 W on a different bike, on a different day, in a different country.
Nothing in running works this way. There is no crank, no rigid transmission, and no sensor in the shoe measuring the force you put into the ground. Instead a device observes motion — accelerations, cadence, ground contact time, vertical oscillation, speed, sometimes grade — and runs a model that outputs a number labeled watts.
There is also no agreed definition of what running power should even include. Should it count the energy stored and returned by your tendons? The work done against air resistance? The cost of swinging your limbs? Different developers answered differently, and those answers are proprietary. This is the reason running power values are not comparable across platforms, and it is not a bug anyone can fix with better hardware.
The two families of implementation
Foot-mounted inertial sensors. A pod on the shoe sits close to the point where force is actually transmitted and samples accelerations at high rate. It also measures ground contact time and leg spring stiffness directly rather than inferring them from the wrist. Stryd is the widely used example.
Watch-native estimates. The watch, sometimes with a chest strap contributing running dynamics, models power from pace, grade, cadence and body mass. No extra hardware, and no sensor anywhere near your foot. Garmin, COROS and Polar all publish something along these lines.
Both are models. The foot pod has better inputs; that is a reason to expect it to be better, not a demonstration that it is.
What has actually been measured
The foot pod against laboratory instruments
Imbach et al. (2020) in Sports ran an incremental test from 8 to 20 km/h on a 200 m indoor track with six recreational runners, against 500 Hz force platforms, a Cosmed K4b2 metabolic analyzer and 400 Hz motion capture. That is a genuine laboratory comparison, and it is the strongest validation this category has.
- Stryd power output against oxygen consumption: R² = 0.82.
- Stryd power output against externally measured mechanical power: R² = 0.88.
- Ground contact time and leg spring stiffness showed no significant difference from the laboratory reference — the underlying biomechanical measurements are good.
Those are respectable numbers and they support the practical claim that the metric tracks effort. But the same paper reported a proportional error in power output that grew with speed, and a systematic underestimation relative to the reference. The device is a better proxy at some speeds than others, which matters if you intend to pace by it across a range of intensities.
The authors also computed the net mechanical efficiency implied by their data at 55 ± 3%. We report that as a flag rather than a finding: an efficiency figure that high is a signal that the quantity being reported is a model output with its own internal conventions, not the physical mechanical work your body performed.
The study's limits are severe and the authors state them: six participants, a correction function estimated from that same small dataset, no maximal speeds, no slopes, no changes of direction, and proprietary algorithms they could not inspect.
Two devices, one runner, one trail — and 82 watts apart
Berzosa et al. (2024) in Sensors did something more directly useful to a buyer. Five participants ran a 2.5 km trail course with 195 m of climbing, twice, a week apart, wearing a Stryd foot pod and a Garmin fēnix 7S Solar with an HRM-PRO simultaneously.
Each device repeated itself almost perfectly. Week to week, Stryd returned intraclass correlations of 0.980 to 0.997 with coefficients of variation of 0.4-2.8%. Garmin returned 0.927 to 0.997 with 0.3-3.3%. Both are excellent. If you use one device and look for changes over time, that is the property you need, and both have it.
The two devices did not agree with each other.
| Metric | Agreement (ICC) | Mean difference | Limits of agreement |
|---|---|---|---|
| Power, whole course | 0.534 (poor) | 82.2 W | -24.9 to 189.3 W |
| Power, uphill | 0.444 (poor) | 109.1 W | 31.4 to 186.9 W |
| Power, downhill | 0.734 | 37.0 W | -49.0 to 123.0 W |
| Speed | 0.991 (excellent) | 0.1 km/h | -1.1 to 1.4 km/h |
| Cadence | 0.969 (good) | 0.1 steps/min | -1.6 to 1.8 steps/min |
| Ground contact time | 0.950 (excellent) | -3.2 ms | -36.7 to 30.3 ms |
Read that table twice. On the things both devices genuinely measure — speed, cadence, contact time — they agree almost exactly. On the thing both devices model — power — they diverge by an average of 82 watts on the same runner on the same run, and the disagreement is worst uphill, where a power number is most useful for pacing.
A crucial caveat, which the authors state plainly: this study had no gold standard. It cannot tell you which device is closer to the truth. It can only tell you that at least one of them is substantially wrong, and that both are wrong in a perfectly consistent way.
Reliable is not the same as valid, and this is the clearest example on the site
Reliability is whether a device gives you the same answer under the same conditions. Validity is whether the answer is correct. They are independent properties, and a device can have the first without the second — a scale that reads eight pounds heavy is perfectly reliable and completely invalid.
Running power is that scale. Intraclass correlations of 0.99 week to week, and 82 watts of disagreement between two implementations of the same concept. That is a reliable measurement of something, and nobody has established that the something is the quantity printed on the screen.
It is worth being fair about what follows from this. Reliability alone is genuinely sufficient for the main thing runners use power for: pacing consistently within a session and comparing efforts on the same device. It is not sufficient for treating your watts as a physiological constant you can carry between devices, publish, or compare with a training partner.
The right way to use it
As a within-device pacing signal. Its real advantage over pace is that it responds to grade and wind immediately, where pace requires you to translate manually and heart rate lags by thirty seconds or more. On a hilly course, that responsiveness is genuinely valuable and the poor cross-device agreement does not diminish it.
Set your zones from your own testing, on your own device. Never import zones or target watts from another platform, another athlete or a coaching plan written around different hardware. Given the measured gap, a target from the wrong device could be off by more than the width of a training zone.
Do not change devices mid-plan. The step change will look like a fitness change and will not be one. If you must switch, retest and rebuild your zones from scratch.
Be most careful uphill. That is where the two implementations diverged most — 109 W apart on average — and, unhelpfully, where runners most want the number.
Do not treat it as more objective than heart rate. Heart rate is measured, with a documented error. Power is modeled, with an undocumented one. Neither is the truth, and the metric that looks most like physics here is the one furthest from a direct measurement.
The record
What has actually been measured
| Device | Metric | What was actually measured | Evidence |
|---|---|---|---|
| Stryd foot pod | Power output vs oxygen consumption | R² = 0.82 across 8-20 km/h; vs externally measured mechanical power R² = 0.88, with proportional error increasing with speed (n = 6) [1] | Peer-reviewed validation |
| Stryd foot pod | Ground contact time and leg spring stiffness | No significant difference from force platforms and motion capture — the biomechanical inputs validated better than the power output built on them [1] | Peer-reviewed validation |
| Stryd vs Garmin fēnix 7S (trail) | Power agreement between devices | ICC 0.534, mean difference 82.2 W, limits of agreement -24.9 to 189.3 W; uphill ICC 0.444 and 109.1 W apart. No gold standard, so neither device is shown correct [2] | Peer-reviewed validation |
| Stryd and Garmin fēnix 7S (trail) | Week-to-week repeatability | Stryd ICC 0.980-0.997 (CV 0.4-2.8%); Garmin ICC 0.927-0.997 (CV 0.3-3.3%) — both highly reliable while disagreeing with each other [2] | Peer-reviewed validation |
| Stryd vs Garmin fēnix 7S (trail) | Speed, cadence, contact time agreement | ICC 0.991, 0.969 and 0.950 — the directly measured quantities agreed closely, isolating power as the modeled outlier [2] | Peer-reviewed validation |
| COROS and Polar running power | Power output accuracy | No validation study exists that we could find. The published comparisons cover Stryd and Garmin only | No validation study exists |
Questions
Frequently asked
Is running power accurate?
It is repeatable rather than verified. Two leading implementations run simultaneously on the same trail returned intraclass correlations above 0.92 against themselves week to week, and disagreed with each other by an average of 82.2 watts — 109.1 watts uphill. That study had no gold standard, so it establishes disagreement rather than which one is right.
Why is running power different from cycling power?
A bike power meter measures deformation of metal under load and computes work directly. No running device measures force applied to the ground. Running power is inferred from motion using a proprietary model, and there is no agreed definition of what it should include, so the values are not comparable across brands.
Is Stryd more accurate than watch-based running power?
It has the better validation record — a laboratory comparison found its power output tracked oxygen consumption at R² = 0.82 and external mechanical power at R² = 0.88, with six participants. It also sits nearer the point where force is transmitted. That is a reason to expect it to be better, not a head-to-head result showing that it is.
Can I use running power targets from a training plan?
Only if the plan was written for the exact device you are using. The measured gap between two implementations was 82 watts on the same run, which is wider than a training zone. Set your targets from your own testing on your own hardware and retest if you switch.
Should I train by power instead of heart rate?
Power responds instantly to hills and wind where heart rate lags by half a minute or more, which makes it a better real-time pacing signal on varied terrain. But heart rate is measured with a documented error and power is modeled with an undocumented one. Using both, on one device, is the defensible position.
Show your working
Sources
Every one of these was read in full before it was cited. Where we could not access a paper, we do not quote its numbers.
- [1]Peer-reviewed validation2020Imbach, Candau, Chailan & Perrey, “Validity of the Stryd Power Meter in Measuring Running Parameters at Submaximal Speeds”, Sports 8(7):103
DOI 10.3390/sports8070103. n = 6 recreational runners, incremental test on a 200 m indoor track from 8 to 20 km/h, against 500 Hz force platforms, a Cosmed K4b2 metabolic analyzer and 400 Hz motion capture. Authors’ stated limits: six participants, a correction function estimated from that same small dataset, no maximal or supramaximal speeds, no slopes or changes of direction, and proprietary algorithms the authors could not inspect.
- [2]Peer-reviewed validation2024Berzosa, Comeras-Chueca, Bascuas, Gutiérrez & Bataller-Cervero, “Assessing Trail Running Biomechanics: A Comparative Analysis of the Reliability of Stryd and GARMIN RP Wearable Devices”, Sensors 24(11):3570
DOI 10.3390/s24113570. n = 5, a 2.5 km trail course with 195 m of climbing, run twice a week apart, comparing a Stryd foot pod against a Garmin fēnix 7S Solar with an HRM-PRO. IMPORTANT: there was no gold standard in this study — it measures how repeatable each device is and how far the two disagree with each other, and it cannot say which one is closer to the truth. That limitation is the reason we cite it the way we do.
- [3]Peer-reviewed validation2025Železnik Mežan, “Accuracy of wearables for determining the maximal oxygen uptake and lactate threshold: a qualitative systematic review”, Frontiers in Sports and Active Living 7:1707991
DOI 10.3389/fspor.2025.1707991. 13 studies, 273 participants (181 men, 92 women), covering Garmin Forerunner 230 and 245, fēnix 3, 6 and 7, Polar V800, Apple Watch Series 7, foot-worn IMUs including Stryd, and chest-worn GNSS. Authors’ stated limits: screened and extracted by a single reviewer, small per-study samples, none of the studies tested the latest smartwatch models, undisclosed algorithms prevent cross-device comparison, and the pooled population skewed to 18-30 year olds.
Read next
Related
- What training load actually measures
The other modeled number, with no validation at all behind it.
- What a lactate threshold estimate is worth
The other pacing anchor your watch infers rather than measures.
- GPS accuracy by device
Speed and distance were the metrics both power devices agreed on — here is how well they are measured.
- Best running watches
Which watches publish running power, and what else they record.
- What VO2 max on a watch is worth
A modeled metric that has, unusually, been measured against a gas analyzer.