Skip to content
Pulse and Path

The evidence scoreboard

Sleep tracking accuracy, device by device

Three devices have been measured against clinical polysomnography. Here is every number from that study, including the ones the press releases rounded.

By Stephen V.Published July 24, 2026
#ad

We earn a commission when you buy through our Amazon links, at no extra cost to you. It never decides a ranking — the rubric that does is published in full, so you can check us rather than take our word for it. How this works.

The short answer

Against clinical polysomnography, sleep-versus-wake agreement was 91-93% across the Oura Ring Gen3, Fitbit Sense 2 and Apple Watch Series 8. Four-stage agreement was lower — 76.3%, 70.9% and 75.0% respectively — and deep-sleep sensitivity varied most of all, from 79.5% down to 50.5%.

The study, and only the study

Every peer-reviewed number in the table above comes from a single paper: Robbins et al., Sensors 24(20):6532 (2024), from Brigham and Women's Hospital. Thirty-five healthy adults aged 20-50, one night of inpatient polysomnography scored in 30-second epochs to AASM guidelines, wearing an Oura Ring Gen3 on the index finger and a Fitbit Sense 2 and Apple Watch Series 8 on opposite wrists.

One paper is a thin evidence base for an entire product category, and that is the honest state of things rather than an editorial choice.

What the numbers mean in practice

Sleep versus wake is a solved problem

91-93% agreement across all three. Total sleep time from any modern tracker is trustworthy enough to act on.

Cohen's kappa is the number to watch, not agreement

Raw agreement flatters a classifier when one class dominates — and light sleep dominates a normal night. Kappa adjusts for the agreement you would get by chance. The four-stage kappas were 0.65 (Oura), 0.60 (Apple), 0.55 (Fitbit): conventionally read as moderate to substantial. Useful, but not the near-certainty a clean hypnogram implies.

Deep sleep is where the devices diverge

79.5%, 61.7% and 50.5% sensitivity. The Apple Watch identified roughly half the deep sleep that polysomnography scored, and shifted about 43 minutes of it into the light bucket.

Which matters because deep sleep is the number people act on. If your watch tells you that you got 45 minutes of deep sleep and the true figure was 88, the anxiety that follows is an artefact of the measurement rather than a fact about your night.

The 76.3% versus 79% discrepancy, again

Oura's press release describes this study as finding 79% four-stage agreement. The paper says 76.3%.

It does not change the ranking and it is not a scandal. It is a clean, checkable example of why the tier system on this site distinguishes a peer-reviewed paper from a manufacturer claim — the two numbers came from the same study and only one of them is in it.

Limits, from the authors themselves

  • Single night, laboratory rather than home conditions.
  • 35 healthy adults without insomnia — not a clinical population.
  • Comparison only during scheduled sleep episodes rather than across 24 hours, which the authors note may inflate concordance.
  • Device failures: 6 of 35 Apple Watches, 2 of 35 Fitbits.
  • The authors state explicitly that findings do not generalize to other devices or subsequent models.
  • Manufacturers publish neither raw sensor data nor their algorithms, so no deeper analysis of why a device errs is possible.

What we could not read

Two further papers are directly relevant and were inaccessible to us on 24 July 2026: the Svensson multi-night Oura validation in Sleep Medicine, and a 2025 systematic review and meta-analysis in OTO Open. A pooled meta-analysis would be stronger evidence than the single study above, so this is a genuine gap in our record. Both are listed in the sources, marked unread. No figure from either appears anywhere on this site.

The record

What has actually been measured

Every row cites a numbered source below, or states outright that no validation study exists. There is no third state.
DeviceMetricWhat was actually measuredEvidence
Oura Ring Gen3Sleep vs wake92.0% agreement, Cohen's κ = 0.60 (p < 0.001) [1]Peer-reviewed validation
Oura Ring Gen3Four-stage classification76.3% agreement, κ = 0.65 — highest of the three devices tested [1]Peer-reviewed validation
Oura Ring Gen3Per-stage sensitivityLight 78.2%, deep 79.5%, REM 76.0% [1]Peer-reviewed validation
Oura Ring Gen3Nightly summary measuresNo significant difference from PSG on seven of eight measures; overestimated sleep latency by 5 minutes [1]Peer-reviewed validation
Apple Watch Series 8Sleep vs wake93.0% agreement, κ = 0.60 — the highest sleep/wake figure of the three [1]Peer-reviewed validation
Apple Watch Series 8Four-stage classification75.0% agreement, κ = 0.60 [1]Peer-reviewed validation
Apple Watch Series 8Deep sleep sensitivity50.5% — the weakest result of the three devices [1]Peer-reviewed validation
Apple Watch Series 8Nightly summary errorOverestimated light sleep by 45 minutes, underestimated deep sleep by 43 minutes, underestimated wake by 7 minutes [1]Peer-reviewed validation
Fitbit Sense 2Sleep vs wake91.0% agreement, κ = 0.52 [1]Peer-reviewed validation
Fitbit Sense 2Four-stage classification70.9% agreement, κ = 0.55 — the weakest of the three [1]Peer-reviewed validation
Fitbit Sense 2Per-stage sensitivityLight 78.0%, deep 61.7%, REM 67.3% [1]Peer-reviewed validation
Fitbit Sense 2Nightly summary errorOverestimated light sleep by 18 minutes, underestimated deep sleep by 15 minutes [1]Peer-reviewed validation
Oura Ring 4 and Ring 5Sleep stagingNo validation of these generations exists that we could find. The published result belongs to the Gen3.No validation study exists
Garmin (all devices)Sleep stagingNo published polysomnography validation exists that we could find, for any Garmin device.No validation study exists
WHOOPSleep stagingNo published PSG sleep-staging validation that we could find. WHOOP 4.0 heart rate was measured in the 2025 placement study, which is a different question.No validation study exists
RingConn Gen 2Sleep stagingNo independent validation study exists that we could find.No validation study exists
Withings Sleep matSleep stagingNo independent validation study exists that we could find. It also infers heart rate through a mattress, the most indirect measurement principle of any device we cover.No validation study exists

Questions

Frequently asked

How accurate is sleep stage tracking?

Moderately. Four-stage agreement with clinical polysomnography ran 70.9-76.3% across the three devices measured, with Cohen's kappa of 0.55-0.65. Sleep versus wake is much better at 91-93%. Treat the stage breakdown as a rough indication and trends over weeks as more meaningful than any single night.

Which device measures deep sleep best?

Of the three measured, the Oura Ring at 79.5% sensitivity, against 61.7% for Fitbit and 50.5% for the Apple Watch. No other device has a published deep-sleep figure we could find.

Why does my sleep tracker disagree with how I feel?

Partly because it may be wrong — deep-sleep sensitivity as low as 50.5% means a device can miss half of it. And partly because how you feel is affected by things no wearable measures: timing relative to your circadian rhythm, sleep fragmentation, alcohol, illness. The tracker measures one input into how you feel, not the output.

Is Garmin or WHOOP sleep tracking accurate?

Nobody has published a polysomnography validation of either that we could find. That is not a criticism of them; it is an empty row in the record, and we would rather show you the empty row than fill it with an opinion.

Show your working

Sources

Every one of these was read in full before it was cited. Where we could not access a paper, we do not quote its numbers.

  1. [1]Peer-reviewed validation2024
    Robbins et al., “Accuracy of Three Commercial Wearable Devices for Sleep Tracking in Healthy Adults”, Sensors 24(20):6532

    Brigham and Women’s Hospital. n = 35 healthy adults, single-night inpatient polysomnography, testing Oura Ring Gen3, Fitbit Sense 2 and Apple Watch Series 8. Authors’ stated limits: one night, lab conditions, healthy adults without insomnia, comparison only during scheduled sleep, and device failures on 6 Apple Watches and 2 Fitbits.

  2. [2]Peer-reviewed validation2025
    “Impact of Anatomical Placement on the Accuracy of Wearable Heart Rate Monitors During Rest and Various Exercise Intensities”, Sensors (Basel)

    n = 28 (14 male, 14 female). Polar H10 chest strap as the reference. Compared Polar Verity Sense (forearm), Garmin Forerunner 55 (wrist) and three simultaneous Whoop 4.0 units at wrist, forearm and upper arm, across rest, cycling warm-up, burpees and a modified Bruce treadmill protocol.

  3. [3]Peer-reviewed validation2024
    Svensson et al., “Validity and reliability of the Oura Ring Generation 3 with sleep staging algorithm 2.0 compared to multi-night ambulatory polysomnography”, Sleep Medicine

    NOT READ — the publisher returned HTTP 403 to our fetch on 2026-07-24. Reported to cover 96 participants and 421,045 epochs. We quote no figure from it. Readers should also note it is the validation of Oura's own staging algorithm, and that a published comment on it exists.

  4. [4]Peer-reviewed validation2025
    Khan et al., “The Oura Ring Versus Medical-Grade Sleep Studies: A Systematic Review and Meta-Analysis”, OTO Open

    NOT READ — the publisher returned HTTP 403 to our fetch on 2026-07-24. A pooled analysis would be stronger evidence than any single study, so this is a real gap in our record rather than a formality.

Read next