The evidence scoreboard
Sleep tracking accuracy, device by device
Three devices have been measured against clinical polysomnography. Here is every number from that study, including the ones the press releases rounded.
We earn a commission when you buy through our Amazon links, at no extra cost to you. It never decides a ranking — the rubric that does is published in full, so you can check us rather than take our word for it. How this works.
The short answer
Against clinical polysomnography, sleep-versus-wake agreement was 91-93% across the Oura Ring Gen3, Fitbit Sense 2 and Apple Watch Series 8. Four-stage agreement was lower — 76.3%, 70.9% and 75.0% respectively — and deep-sleep sensitivity varied most of all, from 79.5% down to 50.5%.
The study, and only the study
Every peer-reviewed number in the table above comes from a single paper: Robbins et al., Sensors 24(20):6532 (2024), from Brigham and Women's Hospital. Thirty-five healthy adults aged 20-50, one night of inpatient polysomnography scored in 30-second epochs to AASM guidelines, wearing an Oura Ring Gen3 on the index finger and a Fitbit Sense 2 and Apple Watch Series 8 on opposite wrists.
One paper is a thin evidence base for an entire product category, and that is the honest state of things rather than an editorial choice.
What the numbers mean in practice
Sleep versus wake is a solved problem
91-93% agreement across all three. Total sleep time from any modern tracker is trustworthy enough to act on.
Cohen's kappa is the number to watch, not agreement
Raw agreement flatters a classifier when one class dominates — and light sleep dominates a normal night. Kappa adjusts for the agreement you would get by chance. The four-stage kappas were 0.65 (Oura), 0.60 (Apple), 0.55 (Fitbit): conventionally read as moderate to substantial. Useful, but not the near-certainty a clean hypnogram implies.
Deep sleep is where the devices diverge
79.5%, 61.7% and 50.5% sensitivity. The Apple Watch identified roughly half the deep sleep that polysomnography scored, and shifted about 43 minutes of it into the light bucket.
Which matters because deep sleep is the number people act on. If your watch tells you that you got 45 minutes of deep sleep and the true figure was 88, the anxiety that follows is an artefact of the measurement rather than a fact about your night.
The 76.3% versus 79% discrepancy, again
Oura's press release describes this study as finding 79% four-stage agreement. The paper says 76.3%.
It does not change the ranking and it is not a scandal. It is a clean, checkable example of why the tier system on this site distinguishes a peer-reviewed paper from a manufacturer claim — the two numbers came from the same study and only one of them is in it.
Limits, from the authors themselves
- Single night, laboratory rather than home conditions.
- 35 healthy adults without insomnia — not a clinical population.
- Comparison only during scheduled sleep episodes rather than across 24 hours, which the authors note may inflate concordance.
- Device failures: 6 of 35 Apple Watches, 2 of 35 Fitbits.
- The authors state explicitly that findings do not generalize to other devices or subsequent models.
- Manufacturers publish neither raw sensor data nor their algorithms, so no deeper analysis of why a device errs is possible.
What we could not read
Two further papers are directly relevant and were inaccessible to us on 24 July 2026: the Svensson multi-night Oura validation in Sleep Medicine, and a 2025 systematic review and meta-analysis in OTO Open. A pooled meta-analysis would be stronger evidence than the single study above, so this is a genuine gap in our record. Both are listed in the sources, marked unread. No figure from either appears anywhere on this site.
The record
What has actually been measured
| Device | Metric | What was actually measured | Evidence |
|---|---|---|---|
| Oura Ring Gen3 | Sleep vs wake | 92.0% agreement, Cohen's κ = 0.60 (p < 0.001) [1] | Peer-reviewed validation |
| Oura Ring Gen3 | Four-stage classification | 76.3% agreement, κ = 0.65 — highest of the three devices tested [1] | Peer-reviewed validation |
| Oura Ring Gen3 | Per-stage sensitivity | Light 78.2%, deep 79.5%, REM 76.0% [1] | Peer-reviewed validation |
| Oura Ring Gen3 | Nightly summary measures | No significant difference from PSG on seven of eight measures; overestimated sleep latency by 5 minutes [1] | Peer-reviewed validation |
| Apple Watch Series 8 | Sleep vs wake | 93.0% agreement, κ = 0.60 — the highest sleep/wake figure of the three [1] | Peer-reviewed validation |
| Apple Watch Series 8 | Four-stage classification | 75.0% agreement, κ = 0.60 [1] | Peer-reviewed validation |
| Apple Watch Series 8 | Deep sleep sensitivity | 50.5% — the weakest result of the three devices [1] | Peer-reviewed validation |
| Apple Watch Series 8 | Nightly summary error | Overestimated light sleep by 45 minutes, underestimated deep sleep by 43 minutes, underestimated wake by 7 minutes [1] | Peer-reviewed validation |
| Fitbit Sense 2 | Sleep vs wake | 91.0% agreement, κ = 0.52 [1] | Peer-reviewed validation |
| Fitbit Sense 2 | Four-stage classification | 70.9% agreement, κ = 0.55 — the weakest of the three [1] | Peer-reviewed validation |
| Fitbit Sense 2 | Per-stage sensitivity | Light 78.0%, deep 61.7%, REM 67.3% [1] | Peer-reviewed validation |
| Fitbit Sense 2 | Nightly summary error | Overestimated light sleep by 18 minutes, underestimated deep sleep by 15 minutes [1] | Peer-reviewed validation |
| Oura Ring 4 and Ring 5 | Sleep staging | No validation of these generations exists that we could find. The published result belongs to the Gen3. | No validation study exists |
| Garmin (all devices) | Sleep staging | No published polysomnography validation exists that we could find, for any Garmin device. | No validation study exists |
| WHOOP | Sleep staging | No published PSG sleep-staging validation that we could find. WHOOP 4.0 heart rate was measured in the 2025 placement study, which is a different question. | No validation study exists |
| RingConn Gen 2 | Sleep staging | No independent validation study exists that we could find. | No validation study exists |
| Withings Sleep mat | Sleep staging | No independent validation study exists that we could find. It also infers heart rate through a mattress, the most indirect measurement principle of any device we cover. | No validation study exists |
Questions
Frequently asked
How accurate is sleep stage tracking?
Moderately. Four-stage agreement with clinical polysomnography ran 70.9-76.3% across the three devices measured, with Cohen's kappa of 0.55-0.65. Sleep versus wake is much better at 91-93%. Treat the stage breakdown as a rough indication and trends over weeks as more meaningful than any single night.
Which device measures deep sleep best?
Of the three measured, the Oura Ring at 79.5% sensitivity, against 61.7% for Fitbit and 50.5% for the Apple Watch. No other device has a published deep-sleep figure we could find.
Why does my sleep tracker disagree with how I feel?
Partly because it may be wrong — deep-sleep sensitivity as low as 50.5% means a device can miss half of it. And partly because how you feel is affected by things no wearable measures: timing relative to your circadian rhythm, sleep fragmentation, alcohol, illness. The tracker measures one input into how you feel, not the output.
Is Garmin or WHOOP sleep tracking accurate?
Nobody has published a polysomnography validation of either that we could find. That is not a criticism of them; it is an empty row in the record, and we would rather show you the empty row than fill it with an opinion.
Show your working
Sources
Every one of these was read in full before it was cited. Where we could not access a paper, we do not quote its numbers.
- [1]Peer-reviewed validation2024Robbins et al., “Accuracy of Three Commercial Wearable Devices for Sleep Tracking in Healthy Adults”, Sensors 24(20):6532
Brigham and Women’s Hospital. n = 35 healthy adults, single-night inpatient polysomnography, testing Oura Ring Gen3, Fitbit Sense 2 and Apple Watch Series 8. Authors’ stated limits: one night, lab conditions, healthy adults without insomnia, comparison only during scheduled sleep, and device failures on 6 Apple Watches and 2 Fitbits.
- [2]Peer-reviewed validation2025“Impact of Anatomical Placement on the Accuracy of Wearable Heart Rate Monitors During Rest and Various Exercise Intensities”, Sensors (Basel)
n = 28 (14 male, 14 female). Polar H10 chest strap as the reference. Compared Polar Verity Sense (forearm), Garmin Forerunner 55 (wrist) and three simultaneous Whoop 4.0 units at wrist, forearm and upper arm, across rest, cycling warm-up, burpees and a modified Bruce treadmill protocol.
- [3]Peer-reviewed validation2024Svensson et al., “Validity and reliability of the Oura Ring Generation 3 with sleep staging algorithm 2.0 compared to multi-night ambulatory polysomnography”, Sleep Medicine
NOT READ — the publisher returned HTTP 403 to our fetch on 2026-07-24. Reported to cover 96 participants and 421,045 epochs. We quote no figure from it. Readers should also note it is the validation of Oura's own staging algorithm, and that a published comment on it exists.
- [4]Peer-reviewed validation2025Khan et al., “The Oura Ring Versus Medical-Grade Sleep Studies: A Systematic Review and Meta-Analysis”, OTO Open
NOT READ — the publisher returned HTTP 403 to our fetch on 2026-07-24. A pooled analysis would be stronger evidence than any single study, so this is a real gap in our record rather than a formality.
Read next
Related
- Most accurate sleep trackers
What to buy, given this evidence.
- Oura Ring 4
The device the strongest result belongs to — one generation back.
- Fix your sleep data
If your tracker disagrees with how you feel, start here.