Brain HealthResearch PaperPaywall

Five Simple Variables Predict 6-Year Dementia Risk With 87% Accuracy

A multinational ML model using only age, sex, education, daily functioning, and cognition predicts mild cognitive impairment risk across three countries.

Sunday, July 19, 2026 8 views
Published in Int J Med Inform
An elderly man seated across from a doctor at a clinic desk, completing a simple paper cognitive test, with a tablet showing a risk score chart visible nearby

Summary

Researchers built a machine learning model to predict who will develop mild cognitive impairment (MCI) over six years using just five variables anyone can measure without expensive tests: age, sex, education level, ability to perform daily tasks, and a baseline cognitive score. Trained on over 11,000 adults aged 60 and older from Chinese, English, and American health studies, the gradient boosting model achieved 87% accuracy internally and remained strong when tested on independent populations in England and the US. The tool is available as a free interactive web app. This matters because catching cognitive decline early — before full dementia sets in — opens a window for lifestyle and medical interventions that may slow progression, and this model makes early screening feasible even in low-resource primary care settings.

0:00--:--

Detailed Summary

Mild cognitive impairment is a critical inflection point in brain aging — it often precedes dementia, yet it represents a window where intervention may still slow or reverse decline. The challenge has always been identification: most existing prediction tools rely on expensive neuroimaging, blood biomarkers, or specialist assessment, placing them out of reach for routine primary care.

This multinational cohort study set out to build a practical, noninvasive prediction model using data from 11,069 adults aged 60 and older drawn from three longitudinal studies: China's CLHLS, England's ELSA, and the US Health and Retirement Study (HRS). Participants were cognitively intact at baseline, and MCI onset was tracked over six years. Over that period, 2,486 participants developed MCI.

Using SHAP-based feature selection and recursive feature elimination, the researchers winnowed hundreds of potential predictors down to just five: age, sex, education, instrumental activities of daily living (IADLs), and baseline cognitive score. Among nine machine learning algorithms tested, a gradient boosting classifier performed best, achieving an internal AUC of 0.869. Crucially, the model generalized well across populations, with AUCs of 0.786 in the English cohort and 0.745 in the US cohort — respectable performance for a five-variable tool applied across different health systems and cultures.

The model was deployed as a free, interactive web application incorporating SHAP explainability, meaning clinicians and patients can see not just a risk score but which factors are driving it — enabling more personalized conversations about prevention.

Implications are substantial for primary care and population health. A validated, noninvasive, cross-national screening tool could enable earlier referral, lifestyle counseling, and clinical trial recruitment. Limitations include reliance on cohort-specific cognitive assessments that may not be fully harmonized, potential residual confounding, and the fact that this summary is based on the abstract only.

Key Findings

  • A five-variable model (age, sex, education, daily functioning, cognition) predicted 6-year MCI risk with AUC of 0.869 internally.
  • Model validated across US and English cohorts with AUCs of 0.745–0.786, demonstrating cross-national generalizability.
  • No expensive tests or invasive biomarkers required — all five predictors are collectible in a standard primary care visit.
  • Interactive web app with SHAP explanations deployed freely, enabling transparent individual risk assessment.
  • 2,486 of 11,069 older adults developed MCI over 6 years, underscoring the scale of the prevention opportunity.

Methodology

Prospective multinational cohort study using data from CLHLS (China), ELSA (England), and HRS (USA), totaling 11,069 adults aged 60+. The Chinese cohort served as internal training/testing; ELSA and HRS were independent external validation sets. Nine ML algorithms were evaluated; feature selection used SHAP values and recursive feature elimination.

Study Limitations

Cognitive assessments differed across cohorts and may not be fully harmonized, potentially affecting cross-national MCI definitions. The model captures baseline risk but cannot account for dynamic changes in risk factors over time. This summary is based on the abstract only, so full methodological details, covariate handling, and sensitivity analyses cannot be evaluated.

Enjoyed this summary?

Get the latest longevity research delivered to your inbox every week.

Enter your email to subscribe: