{"title":"Performance Analysis of Data Mining Techniques for the Prediction Breast Cancer Risk on Big Data","authors":"Solmaz Sohrabei, Alireza Atashi","doi":"10.30699/FHI.V10I1.296","DOIUrl":null,"url":null,"abstract":"Introduction: Early detection breast cancer Causes it most curable cancer in among other types of cancer, early detection and accurate examination for breast cancer ensures an extended survival rate of the patients. Risk factors are an important parameter in breast cancer has an important effect on breast cancer. Data mining techniques have a growing reputation in the medical field because of high predictive capability and useful classification. These methods can help practitioners to develop tools that allow detecting the early stages of breast cancer.Material and Methods: The database used in this paper is provided by Motamed Cancer Institute, ACECR Tehran, Iran. It contains of 7834 records of breast cancer patients clinical and risk factors data. There were 4008 patients (52.4%) with breast cancers (malignant) and the remaining 3617 patients (47.6%) without breast cancers (benign). Support vector machine, multi-layer perceptron, decision tree, K nearest neighbor, random forest, naïve Bayesian models were developed using 20 fields (risk factor) of the database because database feature was restrictions. Used 10-fold crossover for models evaluate. Ultimately, the comparison of the models was made based on sensitivity, specificity and accuracy indicators.Results: Naïve Bayesian and artificial neural network are better models for the prediction of breast cancer risks. Naïve Bayesian had accuracy of 93%, specificity of 93.32%, sensitivity of 95056%, ROC of 0.95 and artificial neural network had accuracy of 93.23%, specificity of 91.98%, sensitivity of 92.69%, and ROC of 0.8.Conclusion: Strangely the different artificial intelligent calculations utilized in this examination yielded close precision subsequently these techniques could be utilized as option prescient instruments in the bosom malignancy risk considers. The significant prognostic components affecting risk pace of bosom disease distinguished in this investigation, which were approved by risk, are helpful and could be converted into choice help devices in the clinical area.","PeriodicalId":154611,"journal":{"name":"Frontiers in Health Informatics","volume":"9 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2021-07-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Frontiers in Health Informatics","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.30699/FHI.V10I1.296","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 2
Abstract
Introduction: Early detection breast cancer Causes it most curable cancer in among other types of cancer, early detection and accurate examination for breast cancer ensures an extended survival rate of the patients. Risk factors are an important parameter in breast cancer has an important effect on breast cancer. Data mining techniques have a growing reputation in the medical field because of high predictive capability and useful classification. These methods can help practitioners to develop tools that allow detecting the early stages of breast cancer.Material and Methods: The database used in this paper is provided by Motamed Cancer Institute, ACECR Tehran, Iran. It contains of 7834 records of breast cancer patients clinical and risk factors data. There were 4008 patients (52.4%) with breast cancers (malignant) and the remaining 3617 patients (47.6%) without breast cancers (benign). Support vector machine, multi-layer perceptron, decision tree, K nearest neighbor, random forest, naïve Bayesian models were developed using 20 fields (risk factor) of the database because database feature was restrictions. Used 10-fold crossover for models evaluate. Ultimately, the comparison of the models was made based on sensitivity, specificity and accuracy indicators.Results: Naïve Bayesian and artificial neural network are better models for the prediction of breast cancer risks. Naïve Bayesian had accuracy of 93%, specificity of 93.32%, sensitivity of 95056%, ROC of 0.95 and artificial neural network had accuracy of 93.23%, specificity of 91.98%, sensitivity of 92.69%, and ROC of 0.8.Conclusion: Strangely the different artificial intelligent calculations utilized in this examination yielded close precision subsequently these techniques could be utilized as option prescient instruments in the bosom malignancy risk considers. The significant prognostic components affecting risk pace of bosom disease distinguished in this investigation, which were approved by risk, are helpful and could be converted into choice help devices in the clinical area.