For decades, a sleep study has boiled down to a single number: the apnea-hypopnea index, or AHI, which counts how many times an hour a person's breathing is disrupted. A new AI model described in Nature Communications suggests that number has been throwing away most of what an overnight sleep study actually knows about a patient's future health.
The model, developed by researchers at Cleveland Clinic and IBM through their decade-long Discovery Accelerator partnership, was trained on roughly 10,000 in-laboratory polysomnography recordings from the Cleveland Clinic Sleep Signals, Testing, and Reports Linked to Patient Traits (STARLIT) registry. Rather than reducing a night of breathing, brain wave, oxygen, and heart rate data down to a severity count, it analyzes the full signal and sorts patients into five risk tiers.
Twice the Death Risk, Invisible to the Standard Score
The gap between the AI's top and bottom risk tiers was stark: patients placed in the highest-risk group had roughly double the five-year mortality risk of those in the lowest-risk group. That distinction didn't show up on the AHI, the metric sleep clinicians have relied on for years to grade apnea severity and guide treatment decisions.
The model's risk categories also tracked with a person's likelihood of developing heart disease and cognitive decline — outcomes that, in current practice, aren't part of a routine sleep study readout at all.
"For decades we have distilled an overnight sleep study into a handful of summary measures," said Dr. Reena Mehra, the study's senior clinical author, now at the University of Washington. "AI gives us the opportunity to move beyond those summaries and learn from the full richness of sleep physiology."
A Blind Spot the AI Doesn't Share
One of the more striking findings involves sex differences. The AHI has long been known to perform better in men than in women as a predictor of downstream health risk — a gap that has quietly shaped how sleep apnea is diagnosed and treated across genders for years. The new AI model predicted outcomes comparably well for both men and women, suggesting it may be picking up on physiological signals that the AHI simply doesn't capture in female patients.
The researchers didn't stop at the original STARLIT cohort. They tested the model against an independent, nationwide patient dataset and found the risk categories held up, a step that matters because AI health tools frequently perform well on the data they were built with and then falter on new populations.
Why a Single Number Was Always a Blunt Instrument
The AHI was designed as a counting exercise: tally the breathing interruptions, divide by hours of sleep. It says nothing about how deep those oxygen dips were, how the brain's electrical activity responded, how heart rate varied through the night, or how disrupted sleep architecture became. A model built to learn from the raw signal, rather than a hand-picked summary of it, has access to all of that texture — which is likely why it can separate patients whose risk the AHI treats as identical.
That has real consequences. Two patients with the same AHI can currently receive the same treatment recommendation, even though one may be at meaningfully higher risk for a heart attack, dementia, or death within five years. A tool that can tell them apart could eventually reshape who gets flagged for more aggressive monitoring or earlier intervention — not just who gets a CPAP prescription.
What This Means for Patients
This tool is not yet something a sleep clinic can order. It was developed and validated on retrospective data, and prospective studies are needed before an AI-generated risk tier could appear alongside a patient's AHI on a sleep study report. Patients currently being evaluated or treated for sleep apnea should continue to follow their AHI-based diagnosis and treatment plan as their doctor prescribes it.
But the direction is notable. If future studies confirm that AI-based risk stratification outperforms the AHI — particularly for women, whose risk the current standard may underestimate — it could eventually help clinicians identify which sleep apnea patients need urgent cardiovascular or cognitive follow-up long before those conditions become clinically apparent, rather than waiting for a single severity number to look bad enough to act on.