Retinal vein occlusion risk prediction without fundus examination using a no-code machine learning tool for tabular data: a nationwide cross-sectional study from South Korea.
Na Hyeon Yu, Daeun Shin, Ik Hee Ryu, Tae Keun Yoo, Kyungmin Koh
{"title":"Retinal vein occlusion risk prediction without fundus examination using a no-code machine learning tool for tabular data: a nationwide cross-sectional study from South Korea.","authors":"Na Hyeon Yu, Daeun Shin, Ik Hee Ryu, Tae Keun Yoo, Kyungmin Koh","doi":"10.1186/s12911-025-02950-8","DOIUrl":null,"url":null,"abstract":"<p><strong>Background: </strong>Retinal vein occlusion (RVO) is a leading cause of vision loss globally. Routine health check-up data-including demographic information, medical history, and laboratory test results-are commonly utilized in clinical settings for disease risk assessment. This study aimed to develop a machine learning model to predict RVO risk in the general population using such tabular health data, without requiring coding expertise or retinal imaging.</p><p><strong>Methods: </strong>We utilized data from the Korea National Health and Nutrition Examination Surveys (KNHANES) collected between 2017 and 2020 to develop the RVO prediction model, with external validation performed using independent data from KNHANES 2021. Model construction was conducted using Orange Data Mining, an open-source, code-free, component-based tool with a user-friendly interface, and Google Vertex AI. An easy-to-use oversampling function was employed to address class imbalance, enhancing the usability of the workflow. Various machine learning algorithms were trained by incorporating all features from the health check-up data in the development set. The primary outcome was the area under the receiver operating characteristic curve (AUC) for identifying RVO.</p><p><strong>Results: </strong>All machine learning training was completed without the need for coding experience. An artificial neural network (ANN) with a ReLU activation function, developed using Orange Data Mining, demonstrated superior performance, achieving an AUC of 0.856 (95% confidence interval [CI], 0.835-0.875) in internal validation and 0.784 (95% CI, 0.763-0.803) in external validation. The ANN outperformed logistic regression and Google Vertex AI models, though differences were not statistically significant in internal validation. In external validation, the ANN showed a marginally significant improvement over logistic regression (P = 0.044), with no significant difference compared to Google Vertex AI. Key predictive variables included age, household income, and blood pressure-related factors.</p><p><strong>Conclusion: </strong>This study demonstrates the feasibility of developing an accessible, cost-effective RVO risk prediction tool using health check-up data and no-code machine learning platforms. Such a tool has the potential to enhance early detection and preventive strategies in general healthcare settings, thereby improving patient outcomes.</p>","PeriodicalId":9340,"journal":{"name":"BMC Medical Informatics and Decision Making","volume":"25 1","pages":"118"},"PeriodicalIF":3.3000,"publicationDate":"2025-03-07","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"BMC Medical Informatics and Decision Making","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.1186/s12911-025-02950-8","RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"MEDICAL INFORMATICS","Score":null,"Total":0}
引用次数: 0
Abstract
Background: Retinal vein occlusion (RVO) is a leading cause of vision loss globally. Routine health check-up data-including demographic information, medical history, and laboratory test results-are commonly utilized in clinical settings for disease risk assessment. This study aimed to develop a machine learning model to predict RVO risk in the general population using such tabular health data, without requiring coding expertise or retinal imaging.
Methods: We utilized data from the Korea National Health and Nutrition Examination Surveys (KNHANES) collected between 2017 and 2020 to develop the RVO prediction model, with external validation performed using independent data from KNHANES 2021. Model construction was conducted using Orange Data Mining, an open-source, code-free, component-based tool with a user-friendly interface, and Google Vertex AI. An easy-to-use oversampling function was employed to address class imbalance, enhancing the usability of the workflow. Various machine learning algorithms were trained by incorporating all features from the health check-up data in the development set. The primary outcome was the area under the receiver operating characteristic curve (AUC) for identifying RVO.
Results: All machine learning training was completed without the need for coding experience. An artificial neural network (ANN) with a ReLU activation function, developed using Orange Data Mining, demonstrated superior performance, achieving an AUC of 0.856 (95% confidence interval [CI], 0.835-0.875) in internal validation and 0.784 (95% CI, 0.763-0.803) in external validation. The ANN outperformed logistic regression and Google Vertex AI models, though differences were not statistically significant in internal validation. In external validation, the ANN showed a marginally significant improvement over logistic regression (P = 0.044), with no significant difference compared to Google Vertex AI. Key predictive variables included age, household income, and blood pressure-related factors.
Conclusion: This study demonstrates the feasibility of developing an accessible, cost-effective RVO risk prediction tool using health check-up data and no-code machine learning platforms. Such a tool has the potential to enhance early detection and preventive strategies in general healthcare settings, thereby improving patient outcomes.
期刊介绍:
BMC Medical Informatics and Decision Making is an open access journal publishing original peer-reviewed research articles in relation to the design, development, implementation, use, and evaluation of health information technologies and decision-making for human health.