This article may contain affiliate links. We may earn a small commission at no extra cost to you if you make a purchase through these links.
Oura Sued Over Sleep Staging: Can a Ring Really Track REM?
A proposed class action says Oura cannot back its sleep-stage accuracy claims. Here is what lab studies show, and how to read your sleep score.
Can a finger ring really tell REM sleep from deep sleep? Partly. In peer-reviewed tests against polysomnography, the clinical sleep study, Oura rings matched a sleep technician's four-stage scoring on about 76% of 30-second epochs in healthy adults, but on only about 53% in an independent study of sleep-clinic patients. A proposed class action filed on August 20, 2026 argues that gap makes Oura's accuracy marketing misleading. Those allegations are unproven, and Oura disputes them.
The lawsuit matters well beyond one company. Every consumer sleep tracker (rings, watches, mattress pads) sells a nightly breakdown of light, deep and REM sleep, and almost none of them measure brain activity. So the case forces a question the category has put off: what does "accurate" actually mean when a device infers brain states from your pulse? Read the numbers correctly and the answer is useful rather than alarming. A ring is a decent trend tracker and a poor stage-by-stage lie detector, and most marketing blurs the two.
What does the Oura lawsuit actually allege?
The case is Surber v. Oura, Inc., docket 3:26-cv-08686 in the U.S. District Court for the Northern District of California, filed August 20, 2026, according to the federal court docket indexed by CourtListener. California resident Madison Surber brought it as a proposed class action, represented by Clarkson Law Firm, which has published the complaint. No class has been certified and no court has ruled on the merits.
According to TechCrunch's report on the filing, the complaint targets Oura marketing that describes the ring as "built for accuracy" with "unparalleled accuracy," and that cites a 79% accuracy figure and, more recently, a claim of 95% sleep-staging accuracy compared with clinical sleep labs. The core argument, as TechCrunch summarizes it, is that "sleep happens in the brain, not on one's finger": proper staging needs scalp electrodes and eye sensors, and the ring instead relies on AI estimates from proxy signals such as heart rate and movement. MedCity News reports that the complaint leans on a 2025 study that found roughly 53% sleep-stage accuracy. The suit seeks an injunction against the challenged advertising plus restitution for buyers.
Oura's answer is on the record. In an August 23, 2026 post, Standing Behind Our Science, the company calls the suit "baseless" and "opportunistic," says it will defend its work "in the appropriate legal forums," and clarifies that its 95% figure refers to telling sleep from wake, not to the four-stage breakdown. Oura also stresses that the ring is a consumer wellness product, not a medical device for diagnosing conditions such as sleep apnea or insomnia.
How is a sleep tracker validated against polysomnography?
Polysomnography (PSG) is the reference standard. Over a night, a lab records brain waves (EEG), eye movements, muscle tone, heart rhythm, breathing and blood oxygen. A trained technician then splits the night into 30-second "epochs" and labels each one: wake, N1, N2, N3 or REM. Consumer trackers collapse N1 and N2 into "light" and call N3 "deep."
A validation study fits the tracker alongside PSG on the same night and compares the two labels epoch by epoch. Three numbers come out of that comparison, and the lawsuit's dispute is largely about which one you quote:
- Accuracy (agreement): the share of epochs where the device and the technician agree. It is easy to understand and easy to flatter, because most of a night is light sleep, so a device that is good at light sleep scores well even if it misreads REM.
- Cohen's kappa: agreement after subtracting what chance alone would produce. It runs from 0 (chance) to 1 (perfect), and it is the harder number to inflate.
- Per-stage sensitivity: of all the epochs the technician scored as, say, REM, the share the device also called REM. This is where individual stages succeed or fail.
Two further details matter. First, two-stage accuracy (sleep versus wake) always looks far better than four-stage accuracy, because it is a much easier task. Second, the reference itself is imperfect: Oura, in an August 28, 2026 post titled Sleep Staging Is Not Guesswork, says two trained human scorers agree on around 83% of epochs in healthy adults, and less in people with sleep disorders. That puts a ceiling on how high any device can realistically score.
The medical establishment's position predates this fight. The American Academy of Sleep Medicine's 2018 position statement on consumer sleep technology concluded that, without validation and FDA clearance, these devices "cannot be utilized for the diagnosis and/or treatment of sleep disorders," while allowing that they may "enhance the patient-clinician interaction." It also called for "access to raw data and algorithms," which is still largely missing across the industry.
What does the peer-reviewed evidence say about Oura?
The studies disagree less than the headlines suggest. They measure different things in different people. Here is how the main papers compare, as of September 2026:
| Study | Population | Four-stage result | Sleep vs. wake | Funding / notes |
|---|---|---|---|---|
| Robbins et al., Sensors, 2024 (Brigham and Women's Hospital) | 35 healthy adults, one lab night; Oura Gen3, Apple Watch Series 8, Fitbit Sense 2 | Oura 76.3% (kappa 0.65); Apple Watch 75.0% (0.60); Fitbit 70.9% (0.55) | Oura 92%; Apple Watch 93%; Fitbit 91% | Funded by Oura; lead author sits on Oura's medical advisory board |
| Svensson et al., Sleep Medicine, 2024 (University of Tokyo) | 96 generally healthy Japanese adults, up to three ambulatory PSG nights each, 421,045 epochs | No significant difference from PSG in light or deep sleep time; REM underestimated by 4.1 to 5.6 minutes | 91.7% to 91.8% | Oura Gen3, algorithm 2.0 |
| Herberger et al., Scientific Reports, 2025 (Charité Berlin) | Sleep-lab patients with sleep and medical disorders; 45 enrolled, 31 usable Oura nights | Oura 53.18% (kappa 0.31); SleepOn 50.48%; Circul 35.06% | Oura 85.03% (kappa 0.43) | No specific funding, no competing interests declared; data from 2022 |
| Khan et al., OTO Open, 2025 (meta-analysis) | 6 studies, 388 participants | No statistically significant pooled difference in light, deep or REM minutes | No significant pooled difference in total sleep time | Mostly healthy adults |
| Jin et al., Journal of Translational Medicine, 2026 (meta-analysis, all finger-worn devices) | 28 studies, 2,426 participants, 11 devices | Kappa ranged from 0.06 to 0.70 across studies | Pooled accuracy 0.87 | Some risk of bias in 27 of 28 studies; only one rated low risk |
Oura's own foundational number, 79% four-stage accuracy, comes from a 2021 paper in Sensors by Altini and Kinnunen. It analyzed 440 nights from 106 people and reported 79% accuracy for four stages and 96% for sleep versus wake when heart-rate-derived and circadian features were included. That work became the basis of Oura's second-generation staging algorithm. In its 2022 announcement, Oura said its paired ring-and-PSG dataset had since grown past 1,200 nights, and it compared the 79% figure with "60-65% agreement" for most wrist-based wearables. The Brigham study, run in healthy adults, puts Oura's four-stage figure a few points lower, and roughly level with an Apple Watch.
Why can Oura be 79% accurate and 53% accurate at the same time?
Both figures can be true because they answer different questions. Four variables explain almost all of the spread:
- Who is sleeping. Healthy volunteers sleep in textbook cycles. Sleep-clinic patients have fragmented nights, apnea events and medications that change heart rhythm, and those are exactly the signals the ring reads. The Charité authors concluded that group averages can look fine while "substantial individual-level inaccuracies" rule rings out for clinical use.
- Which device and software. The Charité data were collected from August to November 2022 on a Gen3 ring. Oura argues the study predates a major algorithm update and the Ring 4, and that it used only two ring sizes, leaving 31% of participants with no usable data. The paper itself confirms the 2022 collection window and sizes 8 and 12.
- How epochs are compared. The Charité team worked from Oura hypnograms exported in 5-minute epochs against 30-second PSG epochs. Oura says that mismatch "artificially" lowers agreement. It is a fair methodological criticism, and it also shows how little raw data consumers and researchers can get out of these devices.
- Who paid. The most favorable head-to-head study was funded by Oura, and its lead author advises the company. That does not make it wrong. Industry-funded validation is normal and the methods are published. But it is the reason independent replication in clinical populations carries extra weight.
That leaves the 95% claim, which is the easiest to critique. Oura now says it refers to sleep-versus-wake detection. Sleep versus wake is a legitimate metric, but a reader who sees "95% Sleep Staging Accuracy" next to a hypnogram of REM and deep sleep will reasonably assume it describes the stages. Whether that framing is legally deceptive is for the court. As a communication choice, it put the easy task's headline number on the hard task's product, and that is the core risk this lawsuit exposes for the whole category.
How should you read your sleep score now?
The evidence supports a clear hierarchy of trust. It applies to Oura and, based on the same body of research, to its ring and watch rivals covered in our sleep tracker shootout and smart ring comparison:
- Most reliable: when you slept and for how long. Sleep-versus-wake agreement sits in the mid-80s to low-90s percent range across the studies above, and pooled total-sleep-time differences are small. Bedtime consistency and total sleep are the numbers to act on.
- Useful as trends: heart rate, HRV and temperature overnight. These are direct physiological measurements rather than inferred brain states. Our HRV accuracy breakdown covers how far to trust them.
- Least reliable: minutes of REM and deep sleep on any single night. Per-stage sensitivity varies widely by study, device and person. Treat one night's "only 40 minutes of REM" as noise. Look at multi-week averages, and only compare nights recorded by the same device.
- Composite scores sit on top of all of the above. A readiness or sleep score blends solid inputs with shaky ones, weighted by a proprietary formula. It is a nudge, not a measurement.
None of this is medical guidance. Consumer trackers are not cleared to diagnose sleep disorders, and the AASM is explicit on that point. If the data and how you feel don't match, particularly with loud snoring, gasping or persistent daytime sleepiness, the conversation belongs with a clinician, who may order an actual sleep study.
What happens next, and why it matters beyond Oura?
The case is at its earliest stage. Filing a complaint establishes nothing about liability. The allegations have to survive Oura's legal challenges before any class is certified, and many consumer class actions end in dismissal or settlement long before a court weighs scientific evidence. Anyone buying or wearing an Oura ring today does not need to act on the lawsuit.
The broader signal is about marketing discipline. Wearable makers have competed on a single headline accuracy percentage without saying which metric it is, which population was tested, or which firmware was used. This complaint shows that plaintiffs' lawyers can now pull those caveats out of the companies' own published studies. The likely industry response, whatever happens in court, is more precise claims: four-stage kappa next to two-stage accuracy, the population, and the study's funder. For buyers, that would be a better outcome than any single verdict.
What to watch: Oura's formal response in the Northern District of California, any independent validation of the Ring 4 and current algorithm in clinical populations, and whether rivals quietly revise their own accuracy language.
Frequently Asked Questions
Is the Oura lawsuit proven?
No. Surber v. Oura, Inc. is a proposed class action filed August 20, 2026 in the U.S. District Court for the Northern District of California. It contains allegations only. No class has been certified and no court has ruled on whether Oura's advertising was misleading. Oura disputes the claims, calls the suit baseless, and says it will defend its work in the appropriate legal forums.
Can a smart ring detect REM sleep without EEG?
It can estimate REM, not measure it. Rings infer stages from heart rate, heart rate variability, breathing rate, temperature and movement, which shift in characteristic ways across sleep stages. In a 2024 study of healthy adults, Oura's four-stage agreement with polysomnography was 76.3%. In a 2025 study of sleep-clinic patients it was 53.18%, with much weaker individual-night accuracy.
What does Oura's 95% accuracy claim refer to?
According to Oura's August 2026 explanation, the 95% figure refers to distinguishing sleep from wake, not to classifying light, deep and REM sleep. Sleep-versus-wake is a much easier task: published studies put Oura in roughly the 85% to 92% range on it, depending on the population. Four-stage agreement is lower in every study, and that is the heart of the complaint.
Is Oura less accurate than an Apple Watch for sleep?
Not on current evidence. In the 2024 Brigham and Women's Hospital study, funded by Oura, the ring scored a four-stage kappa of 0.65 versus 0.60 for Apple Watch Series 8 and 0.55 for Fitbit Sense 2. The Apple Watch underestimated deep sleep by 43 minutes. That was a small, single-night study of healthy adults, so treat it as indicative, not definitive.
Should I stop trusting my sleep tracker?
Trust it for the right things. Total sleep time, bedtime consistency and overnight heart-rate trends are reasonably reliable. Nightly minutes of REM and deep sleep are rough estimates that are best viewed as multi-week averages. No consumer tracker is cleared to diagnose sleep disorders. If symptoms persist, a clinician can order a proper sleep study.
Enjoying this article?
Get more strategic intelligence delivered to your inbox weekly.
Enjoyed this article?
VentureBeast.Tech is independent and reader-supported. If this saved you time, you can buy us a coffee — it keeps the research deep and the site ad-light.
Support us on Ko-fi

Comments (0)
No comments yet. Be the first to share your thoughts!