Dreem Open Datasets: Multi-Scored Sleep Datasets to compare Human and Automated sleep staging
Sleep stage classification constitutes an important element of sleep disorder\ndiagnosis. It relies on the visual inspection of polysomnography records by\ntrained sleep technologists. Automated approaches have been designed to\nalleviate this resource-intensive task. However, such approaches are usually\ncompared to a single human scorer annotation despite an inter-rater agreement\nof about 85 % only. The present study introduces two publicly-available\ndatasets, DOD-H including 25 healthy volunteers and DOD-O including 55 patients\nsuffering from obstructive sleep apnea (OSA). Both datasets have been scored by\n5 sleep technologists from different sleep centers. We developed a framework to\ncompare automated approaches to a consensus of multiple human scorers. Using\nthis framework, we benchmarked and compared the main literature approaches. We\nalso developed and benchmarked a new deep learning method, SimpleSleepNet,\ninspired by current state-of-the-art. We demonstrated that many methods can\nreach human-level performance on both datasets. SimpleSleepNet achieved an F1\nof 89.9 % vs 86.8 % on average for human scorers on DOD-H, and an F1 of 88.3 %\nvs 84.8 % on DOD-O. Our study highlights that using state-of-the-art automated\nsleep staging outperforms human scorers performance for healthy volunteers and\npatients suffering from OSA. Consideration could be made to use automated\napproaches in the clinical setting.\n