Skip to content
Pulse and Path

Rings, bands and recovery

How sleep tracking actually works

Clinical sleep staging is defined by electrical activity in your brain. Your ring measures your pulse. Everything confusing about consumer sleep data follows from that single substitution.

By Stephen V.Published August 13, 2026
#ad

We earn a commission when you buy through our Amazon links, at no extra cost to you. It never decides a ranking — the rubric that does is published in full, so you can check us rather than take our word for it. How this works.

The short answer

Sleep stages are clinically defined by brain activity, eye movement and muscle tone. No consumer wearable measures any of the three. Devices infer stages from heart rate, movement and temperature. Against clinical polysomnography, sleep-versus-wake agreement ran 91-93%, but four-stage classification only 70.9-76.3%.

What the clinical standard actually measures

Polysomnography is the reference against which every sleep device is judged. It is not a better version of what your ring does — it is a different measurement entirely.

A sleep technician attaches electrodes and records at minimum three signals: an electroencephalogram measuring electrical activity at the scalp, an electrooculogram measuring eye movement, and an electromyogram measuring muscle tone, usually at the chin. Sleep stages are then scored in 30-second epochs against published criteria.

Those criteria are defined in terms of those signals. Deep sleep — N3 — is identified by slow delta waves occupying a threshold share of the epoch. REM is identified by a distinctive low-amplitude EEG pattern combined with rapid eye movements and near-total skeletal muscle atonia. The stages are not abstract states that happen to show up on an EEG. They are categories constructed from EEG, EOG and EMG.

That matters more than it first appears. When a wearable reports "deep sleep", it is not measuring a phenomenon that PSG also measures by another route. It is predicting what a technician scoring an EEG would have written down.

What your wearable actually collects

Strip away the app and a consumer sleep tracker has a short list of raw inputs:

  • Movement, from an accelerometer. The oldest sleep-tracking signal there is. Stillness suggests sleep; movement suggests wake or a transition.
  • Heart rate, from an optical sensor. Green or infrared light into the skin, measuring what scatters back off moving blood.
  • Heart rate variability, derived from the same signal. Autonomic tone shifts across the night, and that shift carries real information about state.
  • Respiratory rate, extracted from the modulation breathing imposes on the pulse waveform.
  • Skin temperature, on devices that have a thermistor. Body temperature has a strong circadian rhythm.
  • Blood oxygen estimates, on devices with a second wavelength.

Not one of these is an EEG signal. The device sees the peripheral consequences of what your brain is doing, several inferential steps removed from the thing being scored.

The inference, in the order it happens

  1. Detect sleep onset. Sustained stillness plus a drop in heart rate below your daytime baseline. This is the easiest step and it is the one devices do well.
  2. Segment the night. Split the record into short windows, conventionally 30 seconds to match PSG scoring epochs.
  3. Extract features per window. Mean heart rate relative to your baseline, HRV, movement counts, respiratory rate, temperature trend, and how each is changing.
  4. Classify. A machine-learning model trained on data from people who wore the device while undergoing PSG maps those features to a stage label.
  5. Smooth. Raw classifier output is noisy, so it is constrained by expected sleep architecture — cycles of roughly 90 minutes, deep sleep concentrated early, REM concentrated late.

Step five is worth pausing on. Some of the structure you see in your hypnogram is coming from the model's prior expectation about how a night is shaped, not from anything measured about your night. That is a reasonable engineering decision and it is also why two devices on the same person can produce two plausible-looking, quite different graphs.

The physiological logic, which is genuinely sound

It would be unfair to present this as guesswork. The inference rests on real relationships.

During deep sleep, parasympathetic activity dominates: heart rate falls to its nightly minimum, HRV rises, breathing is regular, and you are close to motionless. During REM, the autonomic picture becomes far more variable — heart rate and breathing become irregular while skeletal muscles are atonic, so you have variability without movement. Light sleep sits between the two, with more micro-movements and brief arousals.

Those are genuine, well-established physiological signatures. A model that watches heart rate, HRV, respiration and movement is watching things that really do differ by stage. The question is only how reliably the peripheral signature identifies the central state, on an individual, on a given night.

How well it works, measured

Robbins et al. (2024) at Brigham and Women's Hospital ran the comparison directly: 35 healthy adults, a single night of inpatient polysomnography, wearing an Oura Ring Gen3, a Fitbit Sense 2 and an Apple Watch Series 8.

Sleep versus wake was good. 91-93% agreement across all three devices. Your total sleep time is broadly trustworthy — which makes sense, because that classification leans hardest on movement and heart-rate level, the two things a wearable genuinely measures.

Four-stage classification was substantially weaker.76.3% for Oura, 75.0% for Apple, 70.9% for Fitbit, with Cohen's kappa of 0.65, 0.60 and 0.55 respectively.

Deep sleep specifically was the weak point. Sensitivity ranged from 79.5% (Oura) to 61.7% (Fitbit) to 50.5% (Apple Watch). The Apple Watch overestimated light sleep by 45 minutes a night and underestimated deep sleep by 43.

Read those numbers as a picture of the method rather than a scoreboard. The step the devices do well is the one supported by what they actually measure. The step they do badly is the one that requires reconstructing a brain state from a pulse. The full device-by-device table is here.

Why placement changes the answer

The best-performing device in that study was a ring, and there is a mechanical reason for it. A finger is a better optical measurement site than a wrist: the tissue is well perfused, the arteries sit close to the surface, and a ring holds a fixed position rather than sliding on a moving joint.

The same principle shows up in the heart-rate literature. The 2025 Sensors placement study ran three identical sensors at wrist, forearm and upper arm simultaneously and found proximal placements beat the wrist at every intensity tested. Better raw signal produces better features, which produces better classification — the chain is the same one.

It is also why an upper-arm night band exists as a product category: it puts the sensor somewhere better and solves the other problem below.

The failure mode nobody talks about

A large share of complaints that begin "my sleep data is wrong" turn out to be "my sleep data is missing". The watch was on the charger at bedtime.

This is not a trivial point. Every one of these devices produces a baseline — your normal resting heart rate, your normal HRV, your normal sleep duration — and a baseline built from four nights a week is mostly noise. A device with a two-day battery that you charge before bed is a device that will never tell you anything reliable, regardless of how good its classifier is.

Which is why battery life is a sleep-tracking specification, not a convenience feature — and why a 14-day watch can beat a better sensor you take off every night.

What no consumer device can do

Score your stages the way a lab does. That requires electrodes on your scalp. Everything else is prediction, and the best prediction available agrees with the clinical standard about three quarters of the time on four-stage classification.

Diagnose anything. Some devices advertise sleep-apnea-related features. No consumer wearable is a diagnostic instrument, and we found no independent validation of those claims. Persistent unrefreshing sleep, loud snoring or witnessed breathing pauses are reasons to see a doctor, not reasons to buy a different tracker.

Tell you why. A device can show you that last night differed from your normal. It cannot tell you whether that was the late meal, the alcohol, the argument or the training.

How to read your own data sensibly

  • Trust total sleep time. 91-93% sleep-wake agreement is genuinely good and it is the number most worth acting on.
  • Treat the stage breakdown as indicative. Particularly deep sleep, where measured sensitivity ranged from 50.5% to 79.5% between devices.
  • Read weeks, not nights. Every device is far more informative about trend than about any single night, and so is your own subjective impression.
  • Do not switch devices mid-baseline. Different sensor, different classifier, different smoothing. The step change is the device, not you.
  • Wear it every night. The most common cause of useless sleep data is missing sleep data.

Questions

Frequently asked

How do sleep trackers know what sleep stage I am in?

They do not know — they predict. A model trained on people who wore the device during a clinical sleep study maps heart rate, heart rate variability, movement, respiratory rate and skin temperature onto stage labels. No consumer wearable measures the brain activity that actually defines the stages.

Are sleep trackers accurate?

For total sleep time, largely yes — sleep-versus-wake agreement was 91-93% against clinical polysomnography. For the four-stage breakdown, much less so: 70.9% to 76.3% agreement, and deep-sleep sensitivity between 50.5% and 79.5% depending on the device.

Why do two devices give me different sleep stages on the same night?

Because each is running a different proprietary classifier on a different raw signal, then smoothing the output against its own assumptions about how a night should be shaped. Some of the structure in your hypnogram comes from the model's expectations rather than from your night.

Is a ring better than a watch for sleep tracking?

On the published evidence, yes — the ring in the polysomnography study had the highest four-stage agreement and the best deep-sleep sensitivity. A finger is a better optical site than a wrist: better perfused, arteries closer to the surface, and no joint moving underneath the sensor.

Can a sleep tracker detect sleep apnea?

No consumer wearable is a diagnostic instrument, and we found no independent validation of the apnea-related features some devices advertise. If you have symptoms — loud snoring, witnessed breathing pauses, persistent unrefreshing sleep — that is a reason to see a doctor rather than to change devices.

Show your working

Sources

Every one of these was read in full before it was cited. Where we could not access a paper, we do not quote its numbers.

  1. [1]Peer-reviewed validation2024
    Robbins et al., “Accuracy of Three Commercial Wearable Devices for Sleep Tracking in Healthy Adults”, Sensors 24(20):6532

    Brigham and Women’s Hospital. n = 35 healthy adults, single-night inpatient polysomnography, testing Oura Ring Gen3, Fitbit Sense 2 and Apple Watch Series 8. Authors’ stated limits: one night, lab conditions, healthy adults without insomnia, comparison only during scheduled sleep, and device failures on 6 Apple Watches and 2 Fitbits.

  2. [2]Peer-reviewed validation2025
    “Impact of Anatomical Placement on the Accuracy of Wearable Heart Rate Monitors During Rest and Various Exercise Intensities”, Sensors (Basel)

    n = 28 (14 male, 14 female). Polar H10 chest strap as the reference. Compared Polar Verity Sense (forearm), Garmin Forerunner 55 (wrist) and three simultaneous Whoop 4.0 units at wrist, forearm and upper arm, across rest, cycling warm-up, burpees and a modified Bruce treadmill protocol.

  3. [3]Peer-reviewed validation2025
    Dial, Hollander, Vatne, Emerson, Edwards & Hagen, “Validation of nocturnal resting heart rate and heart rate variability in consumer wearables”, Physiological Reports 13(16):e70527

    DOI 10.14814/phy2.70527. n = 13 healthy adults (7 male, 6 female, 33.2 ± 8.6 years) across 536 nights, each wearing an Oura Gen 3, Oura Gen 4, Polar Grit X Pro, Garmin fēnix 6 and Whoop 4.0 simultaneously against a Polar H10 single-lead ECG reference sampled at 1000 Hz. Authors’ stated limits: healthy adults only, no atrial-fibrillation population, proprietary and periodically-updated algorithms, and unequal night counts per device. The Garmin was dropped from the resting-heart-rate comparison because the 30-minute window it uses is not timestamped.

  4. [4]Peer-reviewed validation2024
    Svensson et al., “Validity and reliability of the Oura Ring Generation 3 with sleep staging algorithm 2.0 compared to multi-night ambulatory polysomnography”, Sleep Medicine

    NOT READ — the publisher returned HTTP 403 to our fetch on 2026-07-24. Reported to cover 96 participants and 421,045 epochs. We quote no figure from it. Readers should also note it is the validation of Oura's own staging algorithm, and that a published comment on it exists.

  5. [5]Peer-reviewed validation2025
    Khan et al., “The Oura Ring Versus Medical-Grade Sleep Studies: A Systematic Review and Meta-Analysis”, OTO Open

    NOT READ — the publisher returned HTTP 403 to our fetch on 2026-07-24. A pooled analysis would be stronger evidence than any single study, so this is a real gap in our record rather than a formality.

Read next