Ensemble machine learning models for lung cancer incidence risk prediction in the elderly: a retrospective longitudinal study.

IF 3.4 2区 医学 Q2 ONCOLOGY BMC Cancer Pub Date : 2025-01-22 DOI:10.1186/s12885-025-13562-w
Songjing Chen, Sizhu Wu
{"title":"Ensemble machine learning models for lung cancer incidence risk prediction in the elderly: a retrospective longitudinal study.","authors":"Songjing Chen, Sizhu Wu","doi":"10.1186/s12885-025-13562-w","DOIUrl":null,"url":null,"abstract":"<p><strong>Background: </strong>Identifying high risk factors and predicting lung cancer incidence risk are essential to prevention and intervention of lung cancer for the elderly. We aim to develop lung cancer incidence risk prediction model in the elderly to facilitate early intervention and prevention of lung cancer.</p><p><strong>Methods: </strong>We stratified the population into six subgroups according to age and gender. For each subgroup, random forest, extreme gradient boosting, deep neural networks, support vector machine, multiple logistic regression and deep Q network (DQN) models were developed and validated. Models were trained and tested using samples from 2000 to 2015 and independent external validated through those from 2016 to 2019. The suitable model for lung cancer risk prediction and high risk factors identification was chosen based on internal validation and independent external validation.</p><p><strong>Results: </strong>The DQN model achieved the optimal prediction performance in stratified subgroups, with AUROC ranging from 0.937 to 0.953, recall ranging from 0.932 to 0.943, F<sub>2</sub>-score ranging from 0.929 to 0.946, precision ranging from 0.926 to 0.952, F<sub>1</sub>-score ranging from 0.933 to 0.963 and RMSE ranging from 0.21 to 0.27. SHAP values were supplied for model interpretability. High risk factors of lung cancer incidence were identified in the elderly. Men ≥ 65 carrying C > A/G > T mutation had the highest lung cancer incidence decrease of 39.5% after five years quitting in stratified elderly groups, which were 1.83 times more than women ≥ 65 not carrying C > A/G > T mutation.</p><p><strong>Conclusions: </strong>The DQN model may be suitable for identifying high risk factors and predicting lung cancer risk with high performance. The proposed intervention and diagnosis pathways could be used for early screening and intervention before the occurrence of lung cancer, which could help oncologists develop targeted intervention strategies for the stratified elderly to reduce lung cancer incidence and improve therapeutic effect. Proposed method could also be used in predicting the risk of other chronic diseases to help conduct intervention and reduce incidence.</p>","PeriodicalId":9131,"journal":{"name":"BMC Cancer","volume":"25 1","pages":"126"},"PeriodicalIF":3.4000,"publicationDate":"2025-01-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"BMC Cancer","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.1186/s12885-025-13562-w","RegionNum":2,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"ONCOLOGY","Score":null,"Total":0}
引用次数: 0

Abstract

Background: Identifying high risk factors and predicting lung cancer incidence risk are essential to prevention and intervention of lung cancer for the elderly. We aim to develop lung cancer incidence risk prediction model in the elderly to facilitate early intervention and prevention of lung cancer.

Methods: We stratified the population into six subgroups according to age and gender. For each subgroup, random forest, extreme gradient boosting, deep neural networks, support vector machine, multiple logistic regression and deep Q network (DQN) models were developed and validated. Models were trained and tested using samples from 2000 to 2015 and independent external validated through those from 2016 to 2019. The suitable model for lung cancer risk prediction and high risk factors identification was chosen based on internal validation and independent external validation.

Results: The DQN model achieved the optimal prediction performance in stratified subgroups, with AUROC ranging from 0.937 to 0.953, recall ranging from 0.932 to 0.943, F2-score ranging from 0.929 to 0.946, precision ranging from 0.926 to 0.952, F1-score ranging from 0.933 to 0.963 and RMSE ranging from 0.21 to 0.27. SHAP values were supplied for model interpretability. High risk factors of lung cancer incidence were identified in the elderly. Men ≥ 65 carrying C > A/G > T mutation had the highest lung cancer incidence decrease of 39.5% after five years quitting in stratified elderly groups, which were 1.83 times more than women ≥ 65 not carrying C > A/G > T mutation.

Conclusions: The DQN model may be suitable for identifying high risk factors and predicting lung cancer risk with high performance. The proposed intervention and diagnosis pathways could be used for early screening and intervention before the occurrence of lung cancer, which could help oncologists develop targeted intervention strategies for the stratified elderly to reduce lung cancer incidence and improve therapeutic effect. Proposed method could also be used in predicting the risk of other chronic diseases to help conduct intervention and reduce incidence.

查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
求助全文
约1分钟内获得全文 去求助
来源期刊
BMC Cancer
BMC Cancer 医学-肿瘤学
CiteScore
6.00
自引率
2.60%
发文量
1204
审稿时长
6.8 months
期刊介绍: BMC Cancer is an open access, peer-reviewed journal that considers articles on all aspects of cancer research, including the pathophysiology, prevention, diagnosis and treatment of cancers. The journal welcomes submissions concerning molecular and cellular biology, genetics, epidemiology, and clinical trials.
期刊最新文献
Surgery lengthens survival for collecting duct carcinoma: analysis of hospital-based cancer registry data in Japan. Integrating BRCA testing into routine prostate cancer care: a multidisciplinary approach by SIUrO and other Italian Scientific Societies. Improved diagnosis of small cervical lymph node metastasis using postvascular phase perfluorobutane CEUS in cancer patients. The impact of educational intervention based on the theory of planned behavior on preventive behaviors for gastric cancer in obese and smoking individuals. The impact of oxidative balance on all-cause and cause-specific mortality in US adults and cancer survivors: evidence from NHANES 2001-2018.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1