{"title":"[Contribution of the large-scale population cohort in disease risk prediction model study: taking United Kingdom Biobank as an example].","authors":"C X Zhu, Y X Song, Y T Hao, F Chen, Y Y Wei","doi":"10.3760/cma.j.cn112338-20240507-00245","DOIUrl":null,"url":null,"abstract":"<p><p>The disease risk prediction model is the basis of precision prevention and an essential reference for clinical treatment decisions. The development of risk prediction models requires the support of a large amount of high-quality data. A large population cohort study is an important basis for this study. The United Kingdom Biobank (UKB), as a mega-population cohort and biobank, has played an essential role in the exploration of disease etiology and research related to disease prevention and control, with its rich baseline and follow-up data and concepts and mechanisms shared globally. This study followed PRISMA guidelines and included 210 articles with corresponding authors from 18 countries, of which 58 (27.62%) were from the UKB. A total of 491 disease risk prediction models were extracted for cancer, cardiovascular and cerebrovascular diseases, endocrine and metabolic diseases, respiratory diseases, and other diseases and their subgroups, of which 132 were developed by UKB without validation, 183 were developed by UKB with internal validation, 17 were developed by UKB with external validation, and 159 were developed by external development with UKB validation. A total of 188 models used only macro variables (38.29%), and 303 models combined macro and micro variables (61.71%). Model construction methods included survival outcome models, logistic regression, and machine learning. Survival outcome models were dominated by Cox proportional risk regression models and a few models considering competitive risk, accelerated failure models, or different baseline risk functions. Machine learning models included random forest, XGBoost, CatBoost, support vector machine, convolutional neural network, and other methods. The UKB is an essential resource for multiple disease risk prediction modeling studies.</p>","PeriodicalId":23968,"journal":{"name":"中华流行病学杂志","volume":"45 10","pages":"1433-1440"},"PeriodicalIF":0.0000,"publicationDate":"2024-10-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"中华流行病学杂志","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.3760/cma.j.cn112338-20240507-00245","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"Medicine","Score":null,"Total":0}
引用次数: 0
Abstract
The disease risk prediction model is the basis of precision prevention and an essential reference for clinical treatment decisions. The development of risk prediction models requires the support of a large amount of high-quality data. A large population cohort study is an important basis for this study. The United Kingdom Biobank (UKB), as a mega-population cohort and biobank, has played an essential role in the exploration of disease etiology and research related to disease prevention and control, with its rich baseline and follow-up data and concepts and mechanisms shared globally. This study followed PRISMA guidelines and included 210 articles with corresponding authors from 18 countries, of which 58 (27.62%) were from the UKB. A total of 491 disease risk prediction models were extracted for cancer, cardiovascular and cerebrovascular diseases, endocrine and metabolic diseases, respiratory diseases, and other diseases and their subgroups, of which 132 were developed by UKB without validation, 183 were developed by UKB with internal validation, 17 were developed by UKB with external validation, and 159 were developed by external development with UKB validation. A total of 188 models used only macro variables (38.29%), and 303 models combined macro and micro variables (61.71%). Model construction methods included survival outcome models, logistic regression, and machine learning. Survival outcome models were dominated by Cox proportional risk regression models and a few models considering competitive risk, accelerated failure models, or different baseline risk functions. Machine learning models included random forest, XGBoost, CatBoost, support vector machine, convolutional neural network, and other methods. The UKB is an essential resource for multiple disease risk prediction modeling studies.
期刊介绍:
Chinese Journal of Epidemiology, established in 1981, is an advanced academic periodical in epidemiology and related disciplines in China, which, according to the principle of integrating theory with practice, mainly reports the major progress in epidemiological research. The columns of the journal include commentary, expert forum, original article, field investigation, disease surveillance, laboratory research, clinical epidemiology, basic theory or method and review, etc.
The journal is included by more than ten major biomedical databases and index systems worldwide, such as been indexed in Scopus, PubMed/MEDLINE, PubMed Central (PMC), Europe PubMed Central, Embase, Chemical Abstract, Chinese Science and Technology Paper and Citation Database (CSTPCD), Chinese core journal essentials overview, Chinese Science Citation Database (CSCD) core database, Chinese Biological Medical Disc (CBMdisc), and Chinese Medical Citation Index (CMCI), etc. It is one of the core academic journals and carefully selected core journals in preventive and basic medicine in China.