why choose us

300×250 Ad Slot

Research Article: Machine learning based multimodel prediction of poor wellbeing, generalized anxiety disorder, and insomnia among adults in Indian urban informal settlements: toward explainable digital mental health screening

Date Published: 2026-09-30

Abstract:
Mental health morbidity among urban informal settlements-dwelling populations is shaped by a complex interplay of behavioural exposures including screen time, working hours, and sleep duration and sociodemographic vulnerability. We analyzed cross-sectional data from 1,600 adults recruited from 16 clusters within the urban informal settlement field practice area of a medical college's Urban Health Training Centre in Pune, India, to develop and compare ten supervised machine learning algorithms predicting poor wellbeing (WHO-5), moderate-to-severe generalized anxiety disorder (GAD-7), and moderate-to-severe insomnia (ISI); internal consistency of all three instruments was confirmed from item-level responses. Because participants were nested within 16 clusters, our primary analysis used leave-one-cluster-out (LOCO) cross-validation, in which each cluster was held out in turn as an independent test set; a conventional participant-level 80/20 split with 10-fold cross-validation was retained as a secondary analysis. Model discrimination, calibration, and clinical utility were assessed using AUC, sensitivity, specificity, predictive values, Brier score, decision curve analysis, and DeLong's test, each with cluster-level bootstrap 95% confidence intervals. SHAP was used to interpret the best-performing model, and feature-attribution concordance was additionally assessed across the Rashomon set of near-equivalent top models. The WHO-5, GAD-7, and ISI showed acceptable-to-excellent internal consistency in this sample (Cronbach's alpha 0.730–0.938; McDonald's omega 0.815–0.953). LightGBM achieved the highest primary (LOCO-CV) AUC for all three outcomes: 0.977 (95% CI 0.970–0.983) for poor wellbeing, 0.987 (95% CI 0.984–0.992) for GAD, and 0.988 (95% CI 0.982–0.993) for insomnia. Several sociodemographic predictors showed unusually strong univariate associations with the outcomes, traced to a repeating between-cluster pattern in the sampling frame; this structure means that even cluster-aware validation may remain optimistic relative to a genuinely external population, a limitation we discuss in detail. SHAP-based rankings were highly concordant (Spearman ? =?0.84–0.95) across the three top-performing, statistically near-equivalent models. A class-weighting sensitivity analysis improved positive predictive value for poor wellbeing from 0.478 to 0.600 with minimal AUC cost. Gradient-boosted tree ensembles, particularly LightGBM, showed strong and internally consistent discrimination for these three outcomes, assessed using reliable instruments; however, given the between-cluster homogeneity identified in this sample, external validation in an independent population is essential before any claim of generalizable clinical utility.

Introduction:
Mental health morbidity among urban informal settlements-dwelling populations is shaped by a complex interplay of behavioural exposures including screen time, working hours, and sleep duration and sociodemographic vulnerability.

Read more

300×250 Ad Slot