Advancing Alzheimer's disease risk prediction: development and validation of a machine learning-based preclinical screening model in a cross-sectional study.
{"title":"Advancing Alzheimer's disease risk prediction: development and validation of a machine learning-based preclinical screening model in a cross-sectional study.","authors":"Bingsheng Wang, Ruihan Xie, Wenhao Qi, Jiani Yao, Yankai Shi, Xiajing Lou, Chaoqun Dong, Xiaohong Zhu, Bing Wang, Danni He, Yanfei Chen, Shihua Cao","doi":"10.1136/bmjopen-2024-092293","DOIUrl":null,"url":null,"abstract":"<p><strong>Objectives: </strong>Alzheimer's disease (AD) poses a significant challenge for individuals aged 65 and older, being the most prevalent form of dementia. Although existing AD risk prediction tools demonstrate high accuracy, their complexity and limited accessibility restrict practical application. This study aimed to develop a convenience, efficient prediction model for AD risk using machine learning techniques.</p><p><strong>Design and setting: </strong>We conducted a cross-sectional study with participants aged 60 and older from the National Alzheimer's Coordinating Center. We selected personal characteristics, clinical data and psychosocial factors as baseline predictors for AD (March 2015 to December 2021). The study utilised Random Forest and Extreme Gradient Boosting (XGBoost) algorithms alongside traditional logistic regression for modelling. An oversampling method was applied to balance the data set.</p><p><strong>Interventions: </strong>This study has no interventions.</p><p><strong>Participants: </strong>The study included 2379 participants, of whom 507 were diagnosed with AD.</p><p><strong>Primary and secondary outcome measures: </strong>Including accuracy, precision, recall, F1 score, etc. RESULTS: 11 variables were critical in the training phase, including educational level, depression, insomnia, age, Body Mass Index (BMI), medication count, gender, stenting, systolic blood pressure (sbp), neurosis and rapid eye movement. The XGBoost model exhibited superior performance compared with other models, achieving area under the curve of 0.915, sensitivity of 76.2% and specificity of 92.9%. The most influential predictors were educational level, total medication count, age, sbp and BMI.</p><p><strong>Conclusions: </strong>The proposed classifier can help guide preclinical screening of AD in the elderly population.</p>","PeriodicalId":9158,"journal":{"name":"BMJ Open","volume":"15 2","pages":"e092293"},"PeriodicalIF":2.4000,"publicationDate":"2025-02-08","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"BMJ Open","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.1136/bmjopen-2024-092293","RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"MEDICINE, GENERAL & INTERNAL","Score":null,"Total":0}
引用次数: 0
Abstract
Objectives: Alzheimer's disease (AD) poses a significant challenge for individuals aged 65 and older, being the most prevalent form of dementia. Although existing AD risk prediction tools demonstrate high accuracy, their complexity and limited accessibility restrict practical application. This study aimed to develop a convenience, efficient prediction model for AD risk using machine learning techniques.
Design and setting: We conducted a cross-sectional study with participants aged 60 and older from the National Alzheimer's Coordinating Center. We selected personal characteristics, clinical data and psychosocial factors as baseline predictors for AD (March 2015 to December 2021). The study utilised Random Forest and Extreme Gradient Boosting (XGBoost) algorithms alongside traditional logistic regression for modelling. An oversampling method was applied to balance the data set.
Interventions: This study has no interventions.
Participants: The study included 2379 participants, of whom 507 were diagnosed with AD.
Primary and secondary outcome measures: Including accuracy, precision, recall, F1 score, etc. RESULTS: 11 variables were critical in the training phase, including educational level, depression, insomnia, age, Body Mass Index (BMI), medication count, gender, stenting, systolic blood pressure (sbp), neurosis and rapid eye movement. The XGBoost model exhibited superior performance compared with other models, achieving area under the curve of 0.915, sensitivity of 76.2% and specificity of 92.9%. The most influential predictors were educational level, total medication count, age, sbp and BMI.
Conclusions: The proposed classifier can help guide preclinical screening of AD in the elderly population.
期刊介绍:
BMJ Open is an online, open access journal, dedicated to publishing medical research from all disciplines and therapeutic areas. The journal publishes all research study types, from study protocols to phase I trials to meta-analyses, including small or specialist studies. Publishing procedures are built around fully open peer review and continuous publication, publishing research online as soon as the article is ready.