The Effect of Resampling on Classifier Performance: an Empirical Study

U. Pujianto, Muhammad Iqbal Akbar, Niendhitta Tamia Lassela, D. Sutaji
{"title":"The Effect of Resampling on Classifier Performance: an Empirical Study","authors":"U. Pujianto, Muhammad Iqbal Akbar, Niendhitta Tamia Lassela, D. Sutaji","doi":"10.17977/um018v5i12022p87-100","DOIUrl":null,"url":null,"abstract":"An imbalanced class on a dataset is a common classification problem. The effect of using imbalanced class datasets can cause a decrease in the performance of the classifier. Resampling is one of the solutions to this problem. This study used 100 datasets from 3 websites: UCI Machine Learning, Kaggle, and OpenML. Each dataset will go through 3 processing stages: the resampling process, the classification process, and the significance testing process between performance evaluation values of the combination of classifier and the resampling using paired t-test. The resampling used in the process is Random Undersampling, Random Oversampling, and SMOTE. The classifier used in the classification process is Naïve Bayes Classifier, Decision Tree, and Neural Network. The classification results in accuracy, precision, recall, and f-measure values are tested using paired t-tests to determine the significance of the classifier's performance from datasets that were not resampled and those that had applied the resampling. The paired t-test is also used to find a combination between the classifier and the resampling that gives significant results. This study obtained two results. The first result is that resampling on imbalanced class datasets can substantially affect the classifier's performance more than the classifier's performance from datasets that are not applied the resampling technique. The second result is that combining the Neural Network Algorithm without the resampling provides significance based on the accuracy value. Combining the Neural Network Algorithm with the SMOTE technique provides significant performance based on the amount of precision, recall, and f-measure.","PeriodicalId":52868,"journal":{"name":"Knowledge Engineering and Data Science","volume":" ","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2022-06-07","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Knowledge Engineering and Data Science","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.17977/um018v5i12022p87-100","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0

Abstract

An imbalanced class on a dataset is a common classification problem. The effect of using imbalanced class datasets can cause a decrease in the performance of the classifier. Resampling is one of the solutions to this problem. This study used 100 datasets from 3 websites: UCI Machine Learning, Kaggle, and OpenML. Each dataset will go through 3 processing stages: the resampling process, the classification process, and the significance testing process between performance evaluation values of the combination of classifier and the resampling using paired t-test. The resampling used in the process is Random Undersampling, Random Oversampling, and SMOTE. The classifier used in the classification process is Naïve Bayes Classifier, Decision Tree, and Neural Network. The classification results in accuracy, precision, recall, and f-measure values are tested using paired t-tests to determine the significance of the classifier's performance from datasets that were not resampled and those that had applied the resampling. The paired t-test is also used to find a combination between the classifier and the resampling that gives significant results. This study obtained two results. The first result is that resampling on imbalanced class datasets can substantially affect the classifier's performance more than the classifier's performance from datasets that are not applied the resampling technique. The second result is that combining the Neural Network Algorithm without the resampling provides significance based on the accuracy value. Combining the Neural Network Algorithm with the SMOTE technique provides significant performance based on the amount of precision, recall, and f-measure.
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
重新采样对分类器性能的影响:一项实证研究
数据集上的不平衡类是一个常见的分类问题。使用不平衡的类数据集会导致分类器性能的下降。重采样是解决这一问题的方法之一。这项研究使用了来自3个网站的100个数据集:UCI机器学习、Kaggle和OpenML。每个数据集将经过3个处理阶段:重采样过程、分类过程、分类器组合的性能评价值与重采样使用配对t检验的显著性检验过程。在此过程中使用的重采样是随机欠采样,随机过采样和SMOTE。在分类过程中使用的分类器是Naïve贝叶斯分类器,决策树和神经网络。分类结果的准确性、精密度、召回率和f测量值使用配对t检验来确定分类器性能的显著性,这些数据集来自未重新采样的数据集和应用重新采样的数据集。配对t检验也用于找到分类器和重采样之间的组合,从而产生显著的结果。这项研究得到了两个结果。第一个结果是,与未应用重采样技术的数据集相比,对不平衡类数据集进行重采样对分类器性能的影响更大。第二个结果是结合不重采样的神经网络算法提供了基于精度值的意义。将神经网络算法与SMOTE技术相结合,基于精度、召回率和f-measure的数量提供了显著的性能。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
自引率
0.00%
发文量
4
审稿时长
8 weeks
期刊最新文献
Optimizing Random Forest Algorithm to Classify Player's Memorisation via In-game Data Long-Term Traffic Prediction Based on Stacked GCN Model Round-Robin Algorithm in Load Balancing for National Data Centers K-Means Clustering and Multilayer Perceptron for Categorizing Student Business Groups Maximum Marginal Relevance and Vector Space Model for Summarizing Students' Final Project Abstracts
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1