Sleep research article
Interpretable machine learning for depression symptom classification in NHANES: Performance in a curated high-confidence corpus and the full survey population.
Authors: Salimi O , Ghahramani Z , Zenouz MT , Heidari M , Amini A , Nouroozi F , Fathi M , Tavasol A , Bateni MR
One-line summary
A sleep science research article on Interpretable machine learning for depression symptom classification in NHANES: Performance in a curated high-confidence corpus and the full survey population..
Sleep health notes
Sleep health notes will be added by the Sleepatch editorial team.
中文解读
中文解读待补充:本站会优先为失眠研究、睡眠质量改善、昼夜节律等高价值睡眠研究添加中文说明。
Original abstract
<h4>Background</h4>Depressive disorders are among the most common psychiatric conditions worldwide and frequently remain undetected or misinterpreted. Scalable computational tools may help characterize depressive-symptom patterns in large health datasets, but their clinical use requires realistic validation and calibration.<h4>Objective</h4>We evaluated machine-learning models for classifying PHQ-9 depressive-symptom status in NHANES and compared performance in a curated high-confidence corpus and the full analytic population.<h4>Methods</h4>This cross-sectional analysis included 37,959 NHANES participants with valid PHQ-9 data and prespecified demographic, sleep, and dietary predictors. The primary outcome was mild-or-greater depressive symptoms, defined as PHQ-9 ≥ 5. A balanced high-confidence corpus was constructed using clear PHQ-9 rules, teacher-model confidence ranking, Isolation Forest filtering, and class-specific resampling. A LightGBM classifier was compared with logistic regression, random forest, gradient boosting, support vector machine, XGBoost, k-nearest neighbors, and Gaussian naïve Bayes. Performance, calibration, temporal validation, sensitivity analyses, and SHAP-based interpretability were assessed.<h4>Results</h4>The curated corpus contained 5000 balanced observations. LightGBM achieved near-perfect curated-holdout performance, but full-population performance was moderate: accuracy 0.678, F1-score 0.503, precision 0.407, recall 0.659, and ROC-AUC 0.726. Comparator models showed similar full-population discrimination. Removing sleep predictors reduced ROC-AUC to 0.611. Full-population calibration was poor, and self-reported trouble sleeping was the dominant contributor.<h4>Conclusion</h4>Interpretable machine learning identified reproducible depressive-symptom patterns in NHANES, largely driven by sleep variables. However, modest precision and poor calibration preclude stand-alone clinical use without external validation and recalibration.
Links and sources
This content is provided for informational and educational purposes only and does not constitute medical advice, diagnosis, or treatment. Sleep disorders, chronic insomnia, sleep apnea, and other conditions must be evaluated and treated by a qualified healthcare professional. If you experience persistent or severe sleep problems, consult a licensed physician or sleep specialist. Research cited refers to peer-reviewed studies; individual results may vary. Sleepatch does not endorse any specific medication, supplement, or therapy.
Want a personalized sleep improvement plan?
Sleepatch can prepare a customized sleep wellness program, insomnia relief guide, and evidence-based sleep coaching based on your needs.
Explore sleep services
Comments