Skip to content
Reference Plate
RP-REP-2026-056 subjects · checked 2026-08-20

Which sleep tracker is most accurate, and can any of them measure sleep stages?

Short answer · checked 2026-08-20

Fitbit is the most accurate consumer tracker for total sleep time in the published record: +2.6 min against polysomnography over 49 nights, and second or third of six devices in two later labs. No consumer tracker measures sleep stages well enough to act on. The best stage agreement ever published for a wrist device is kappa 0.53; Garmin sits at 0.21.

Callum Reith and Nandini Rege · published

What we compared

Listed as tested, not ranked

Subjects in RP-REP-2026-05
#SubjectTypeOfficial site
1Apple WatchDeviceapple.com
2Eight SleepDeviceeightsleep.com
3FitbitDevicefitbit.com
4GarminDevicegarmin.com
5Oura RingDeviceouraring.com
6WhoopDevicewhoop.com

We have the polysomnograph and we have not used it yet

This desk has a polysomnograph. It has not yet published a sleep run with it.

Method clause M-04 — sleep staging and duration against a Type II ambulatory polysomnograph — is drafted. The rig was built and checked on 2026-04-21. No run using it has published. So nothing below is our measurement, and we are not going to imply otherwise.

What follows is the published record, read on 2026-08-10, with every figure attributed to whoever produced it, under clause M-09. Where two labs disagree we say so rather than averaging them into a number that belongs to neither.

One framing point before the tables. The reference standard for sleep measurement is polysomnography: EEG, EOG and EMG, scored in 30-second epochs by a human being. It reads electrical activity from the brain. A wrist tracker reads acceleration and an optical pulse signal from a peripheral artery. It is not a worse version of the same instrument. It is a different instrument inferring the output of the first one. That distinction predicts, correctly, which parts of the output survive validation and which do not.

Total sleep time is the part that works

Figures below are the reported bias against simultaneous polysomnography. Positive means the device recorded more sleep than the reference.

DeviceTST biasnStudyYear
Fitbit Alta HR+2.6 min49 nightsChinoy et al., Sleep 44(5)2021
Whoop 3.0+8.2 min86 sleepsMiller et al., J Sports Sci 38(22)2020
Oura Ring Gen 1no significant bias41 peoplede Zambotti et al., Behav Sleep Med 17(2)2019
Garmin Fenix 5S+43.7 min29 nightsChinoy et al., Sleep 44(5)2021
Garmin Vivosmart 3+46.8 min43 nightsChinoy et al., Sleep 44(5)2021
Pooled, 24 studies−16.9 min798 peopleLee et al., J Clin Sleep Med 21(3)2025

Two things in that table need saying plainly.

First, the spread between brands is larger than the spread between a good tracker and the reference. A Fitbit Alta HR landed within three minutes of a scored polysomnogram. Two Garmin models were three quarters of an hour out, in the same laboratory, on the same protocol, scored the same way. That is not a rounding difference; it is the difference between a usable number and a decorative one.

Second, the pooled meta-analytic figure runs the other way. Lee and colleagues pooled 24 studies and 798 participants and found consumer wrist devices underestimated total sleep time by 16.9 min (95% CI −26.3 to −7.4). Chinoy’s devices mostly overestimated. Both results are real. They cover different device generations and different populations, and the honest reading is that the direction of the error is not stable across the category even though its magnitude is usually tolerable. Under clause M-03 we would call most of these devices tied on total sleep time and Garmin separate, in the wrong direction.

The failure mode is wake, not sleep

This is the mechanism behind almost every complaint readers send us, so it gets its own table. Sensitivity here is the proportion of true sleep epochs the device called sleep. Specificity is the proportion of true wake epochs it called wake.

DeviceSensitivity to sleepSpecificity for wakeSource
Fitbit Alta HR0.950.54Chinoy 2021, n = 34
Whoop 3.00.950.51Miller 2020, n = 12
Oura Ring Gen 10.960.48de Zambotti 2019, n = 41
Garmin Fenix 5S0.990.18Chinoy 2021, n = 34
Garmin Vivosmart 30.990.19Chinoy 2021, n = 34

Every device is excellent at recognising sleep and poor at recognising wake. That asymmetry is not a bug in one product, it is what the sensor set can support: a person lying still with a low resting heart rate is, to an accelerometer and a photoplethysmograph, indistinguishable from a person asleep. Garmin’s 0.99 sensitivity paired with 0.18 specificity is the extreme case — a device that calls almost everything sleep will score near-perfectly on sleep and almost never catch you awake.

The practical consequence: these devices systematically flatter your night. If you lay awake for forty minutes without moving much, most of them will sell it back to you as sleep.

Sleep stages are the part to ignore

Stage classification is where the category should be read as unvalidated. The most useful recent evidence is a 2025 study from Hasselt and Antwerp — 62 adults, one laboratory night each, funded by the Flemish government agency VLAIO, with the authors declaring no financial arrangements or affiliations. That independence matters and we weight it accordingly.

DeviceCohen’s κWakeLightDeepREM
Apple Watch Series 80.5352.2%83.3%50.7%68.6%
Fitbit Sense0.4248.8%73.3%50.9%61.3%
Fitbit Charge 50.4147.7%72.4%51.5%60.0%
Whoop 4.00.3740.1%62.0%69.6%62.0%
Withings Scanwatch0.2231.1%53.0%66.7%
Garmin Vivosmart 40.2127.6%60.3%47.5%33.1%

Schyvens et al., “A performance validation of six commercial wrist-worn wearable sleep-tracking devices for sleep stage scoring compared to polysomnography”, SLEEP Advances 6(2) zpaf021, 2025. Per-stage figures are the percentage of epochs correctly classified.

The best figure in that table, κ = 0.53, is conventionally described as moderate agreement. It is the highest published stage-scoring agreement we found for a wrist-worn consumer device, and it would not be accepted as an equivalent method for anything clinical. Below it, four of six devices fall into the fair range. Garmin’s REM classification at 33.1% is worse than it looks, because a classifier can reach that by guessing with the right base rates.

Deep sleep is the number readers care about most and the number to trust least. In a 2026 in-home study, an Oura Ring Gen 3 overstated deep sleep by 71.5 min in older adults and 81.0 min in younger adults; a Fitbit overstated it by 29.3 min in older adults and understated it by 18.3 min in younger ones. Same night, same reference, opposite direction by brand and by age group.

Two labs, two answers, one watch

Here is the result that should stop anyone quoting a single stage figure as fact.

In the Belgian study above, the Apple Watch Series 8 was the best of six devices at staging, κ = 0.53. In a 2023 two-centre Korean study of 75 people, an Apple Watch 8 was seventh of eleven, κ = 0.298 — below an Oura Ring 3 at 0.349 and a Fitbit Sense 2 at 0.419.

Same product generation. Two competent labs. Kappa 0.53 against 0.298.

Different participants, different scoring teams and different protocols will do that, and neither study is wrong. What it tells you is that stage accuracy for these devices is not a stable property you can look up. It is a property of the device, the sleeper and the night, and a single published kappa is a weak predictor of what any individual will get.

One disclosure attaches to that Korean study, and clause M-09 requires us to carry it: its top-ranked entrant was SleepRoutine, made by Asleep, and the paper discloses that authors affiliated with Asleep hold stock or stock options in Asleep. The study appears carefully done and we use its comparative figures. We do not lean on its winner.

Age moves the error more than brand does

Most validation samples are young, healthy and sleeping in a laboratory. A 2026 study ran four devices at home against polysomnography in 19 adults aged 56 to 80 and 13 aged 19 to 24, and the age split dominated everything else.

DeviceTST bias, 56–80TST bias, 19–24
Fitbit−74.5 min−23.1 min
Oura Ring Gen 3−75.5 min−15.5 min
Withings Sleep Mat−45.7 min−4.8 min
SleepScore Max−56.5 min−14.7 min

Searles et al., SLEEP Advances 7(1) zpag006, published 2026-01-12.

An hour and a quarter of missing sleep is not a tracker being slightly off. Older sleep is more fragmented, with more brief awakenings, and a movement-and-pulse algorithm tuned on young sleepers degrades sharply against it. If you are over 55, the validation literature is largely not about you, and the figures in the first table of this report should be read as a best case.

Eight Sleep: nothing published to report

Eight Sleep is in this report because readers asked about it. In a literature search run on 2026-08-10 we found no peer-reviewed polysomnography validation of the Pod, and it did not appear in any of the multi-device comparisons we reviewed.

We are stating what our search found on a date, not making a claim about the product. A device with no published validation is not thereby inaccurate. It is unverified, which is a different thing and, for a reader deciding where to put several hundred pounds, arguably a more important one. If a validation exists that we missed, send it and we will log the correction.

Device by device

Fitbit — the pick on total sleep time, and the most consistent performer across independent labs: tightest duration figure in Chinoy 2021, second and third of six in Schyvens 2025, and best of the major consumer brands in Lee 2023. It is not good at staging. Nothing is.

Apple Watch — the single best published stage-scoring result for a wrist device, and a much weaker one from a second lab. Best-in-class wake specificity at 52.2%, which is still a coin flip.

Whoop — competent on duration (+8.2 min, though from only 12 participants) and mid-pack on staging at κ = 0.37. Its 69.6% deep-sleep epoch accuracy was the highest of the six in Schyvens, which is worth noting and not worth much on its own.

Oura Ring — no significant duration bias in the original 2019 validation, κ = 0.349 in 2023, and a large deep-sleep overestimate in 2026. Finger photoplethysmography has a cleaner pulse signal than the wrist; it does not appear to have converted that into better staging.

Garmin — the clear warning in this report. Two models 43.7 and 46.8 min long on total sleep time, wake specificity of 0.18 and 0.19, and the lowest staging kappa of six devices at 0.21. Different Garmin models were tested in each study and the results were poor in both.

Eight Sleep — unverified, as above.

What would change our mind

Publishing RP-RUN on sleep, which is the point of clause M-04 existing. Until we run it, this report is a reading of other people’s work and is labelled as such.

Specifically, we would revisit this if: a manufacturer published epoch-level agreement against polysomnography on a pre-declared sample rather than summary statistics; a stage-scoring kappa above 0.60 were replicated by a second unaffiliated lab on its own participants; or a validation sample were published with a median age above 50. None of those exist for any device in this report as of 2026-08-10.

Until then the defensible position is narrow and we will hold it: use these devices for total sleep time and for week-to-week trends in your own data, and do not make decisions from the coloured stage bars.

Questions we got

Sent to the desk · answered in full

Which sleep tracker is most accurate for total sleep time?

Fitbit, on the published evidence as of 2026-08-10. A Fitbit Alta HR came within +2.6 min of polysomnography across 49 nights in a 2021 Naval Health Research Center study, the tightest total-sleep-time figure of the four wearables tested there. Later Fitbit models placed second and third of six devices in a 2025 Belgian lab study. Garmin is the weakest of the major brands on this measure: two Garmin models overestimated total sleep time by 43.7 and 46.8 min in the same 2021 study.

Are sleep stages from a smartwatch accurate?

No, not accurately enough to act on. The best sleep-stage agreement published for a wrist device is Cohen's kappa 0.53, for an Apple Watch Series 8 in a 2025 study of 62 adults, which is conventionally read as moderate agreement. Deep-sleep detection across six devices in that study ranged from 47% to 70% of epochs. An Oura Ring Gen 3 overstated deep sleep by 71 to 81 min per night in a 2026 study. Wrist and finger sensors read movement and pulse, not brain activity, so stage boundaries are inferred rather than measured.

Why does my tracker say I was asleep when I was lying awake?

Because these devices are tuned to detect sleep, not wake, and the two are not symmetrical. Across every validated device, sensitivity to sleep runs above 90% while specificity for wake runs between 18% and 54%. A person lying still with a resting heart rate looks identical to a sleeping person through an accelerometer and an optical pulse sensor. The practical effect is that quiet wakefulness is scored as sleep, so trackers tend to flatter your night.

Is the Oura Ring more accurate than Whoop?

Not on the evidence available on 2026-08-10, and the gap between them is smaller than the gap between studies. Whoop 4.0 scored kappa 0.37 for sleep staging in a 2025 study of 62 adults; an Oura Ring 3 scored kappa 0.349 in a 2023 multicentre study of 75 people. Those were different labs, different participants and different protocols, so the two figures cannot be subtracted from one another. Both sit in the range where stage output should not be treated as a measurement.

Has Eight Sleep been validated against polysomnography?

We could not find a peer-reviewed polysomnography validation of the Eight Sleep Pod in a literature search run on 2026-08-10, and it did not appear in any of the multi-device comparisons we reviewed. That is a statement about the published record on that date, not a measurement of the product. Absence of a published study is not evidence that a device is inaccurate; it means nobody outside the company has shown that it is accurate.

Do sleep trackers work as well for older adults?

No. In a 2026 in-home study, four consumer devices underestimated total sleep time by 45.7 to 75.5 min in adults aged 56 to 80, against 4.8 to 23.1 min in adults aged 19 to 24. Fragmented sleep with more brief awakenings is harder for a movement-and-pulse algorithm to score, and most validation samples are young and healthy. Age shifted the error more than the choice of brand did.

Every product named here links to its own site. We hold no affiliate account with any of them. Procedure: methods. Corrections: the log.