Ashir Javeed, Peter Anderberg, Muhammad Asim Saleem, Ahmad Nauman Ghazi, Johan Sanmartin Berglund
{"title":"Unveiling Cancer: A Data-Driven Approach for Early Identification and Prediction Using F-RUS-RF Model","authors":"Ashir Javeed, Peter Anderberg, Muhammad Asim Saleem, Ahmad Nauman Ghazi, Johan Sanmartin Berglund","doi":"10.1002/ima.23221","DOIUrl":null,"url":null,"abstract":"<p>Globally, cancer is the second-leading cause of death after cardiovascular disease. To improve survival rates, risk factors and cancer predictors must be identified early. From the literature, researchers have developed several kinds of machine learning-based diagnostic systems for early cancer prediction. This study presented a diagnostic system that can identify the risk factors linked to the onset of cancer in order to anticipate cancer early. The newly constructed diagnostic system consists of two modules: the first module relies on a statistical F-score method to rank the variables in the dataset, and the second module deploys the random forest (RF) model for classification. Using a genetic algorithm, the hyperparameters of the RF model were optimized for improved accuracy. A dataset including 10 765 samples with 74 variables per sample was gathered from the Swedish National Study on Aging and Care (SNAC). The acquired dataset has a bias issue due to the extreme imbalance between the classes. In order to address this issue and prevent bias in the newly constructed model, we balanced the classes using a random undersampling strategy. The model's components are integrated into a single unit called F-RUS-RF. With a sensitivity of 92.25% and a specificity of 85.14%, the F-RUS-RF model achieved the highest accuracy of 86.15%, utilizing only six highly ranked variables according to the statistical F-score approach. We can lower the incidence of cancer in the aging population by addressing the risk factors for cancer that the F-RUS-RF model found.</p>","PeriodicalId":14027,"journal":{"name":"International Journal of Imaging Systems and Technology","volume":"34 6","pages":""},"PeriodicalIF":3.0000,"publicationDate":"2024-11-20","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1002/ima.23221","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"International Journal of Imaging Systems and Technology","FirstCategoryId":"94","ListUrlMain":"https://onlinelibrary.wiley.com/doi/10.1002/ima.23221","RegionNum":4,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"ENGINEERING, ELECTRICAL & ELECTRONIC","Score":null,"Total":0}
引用次数: 0
Abstract
Globally, cancer is the second-leading cause of death after cardiovascular disease. To improve survival rates, risk factors and cancer predictors must be identified early. From the literature, researchers have developed several kinds of machine learning-based diagnostic systems for early cancer prediction. This study presented a diagnostic system that can identify the risk factors linked to the onset of cancer in order to anticipate cancer early. The newly constructed diagnostic system consists of two modules: the first module relies on a statistical F-score method to rank the variables in the dataset, and the second module deploys the random forest (RF) model for classification. Using a genetic algorithm, the hyperparameters of the RF model were optimized for improved accuracy. A dataset including 10 765 samples with 74 variables per sample was gathered from the Swedish National Study on Aging and Care (SNAC). The acquired dataset has a bias issue due to the extreme imbalance between the classes. In order to address this issue and prevent bias in the newly constructed model, we balanced the classes using a random undersampling strategy. The model's components are integrated into a single unit called F-RUS-RF. With a sensitivity of 92.25% and a specificity of 85.14%, the F-RUS-RF model achieved the highest accuracy of 86.15%, utilizing only six highly ranked variables according to the statistical F-score approach. We can lower the incidence of cancer in the aging population by addressing the risk factors for cancer that the F-RUS-RF model found.
期刊介绍:
The International Journal of Imaging Systems and Technology (IMA) is a forum for the exchange of ideas and results relevant to imaging systems, including imaging physics and informatics. The journal covers all imaging modalities in humans and animals.
IMA accepts technically sound and scientifically rigorous research in the interdisciplinary field of imaging, including relevant algorithmic research and hardware and software development, and their applications relevant to medical research. The journal provides a platform to publish original research in structural and functional imaging.
The journal is also open to imaging studies of the human body and on animals that describe novel diagnostic imaging and analyses methods. Technical, theoretical, and clinical research in both normal and clinical populations is encouraged. Submissions describing methods, software, databases, replication studies as well as negative results are also considered.
The scope of the journal includes, but is not limited to, the following in the context of biomedical research:
Imaging and neuro-imaging modalities: structural MRI, functional MRI, PET, SPECT, CT, ultrasound, EEG, MEG, NIRS etc.;
Neuromodulation and brain stimulation techniques such as TMS and tDCS;
Software and hardware for imaging, especially related to human and animal health;
Image segmentation in normal and clinical populations;
Pattern analysis and classification using machine learning techniques;
Computational modeling and analysis;
Brain connectivity and connectomics;
Systems-level characterization of brain function;
Neural networks and neurorobotics;
Computer vision, based on human/animal physiology;
Brain-computer interface (BCI) technology;
Big data, databasing and data mining.