作者识别与机器学习算法

International Journal of Multidisciplinary Studies and Innovative Technologies Pub Date : 1900-01-01 DOI:10.36287/ijmsit.6.1.45

İbrahim Yülüce, Feriştah Dalkılıç

{"title":"作者识别与机器学习算法","authors":"İbrahim Yülüce, Feriştah Dalkılıç","doi":"10.36287/ijmsit.6.1.45","DOIUrl":null,"url":null,"abstract":"– Author identification is one of the application areas of text mining. It deals with the automatic prediction of the potential author of an electronic text among predefined author candidates by using author specific writing styles. In this study, we conducted an experiment for the identification of the author of a Turkish language text by using classical machine learning methods including Support Vector Machines (SVM), Gaussian Naive Bayes (GaussianNB), Multi Layer Perceptron (MLP), Logistic Regression (LR), Stochastic Gradient Descent (SGD) and ensemble learning methods including Extremely Randomized Trees (ExtraTrees), and eXtreme Gradient Boosting (XGBoost). The proposed method was applied on three different sizes of author groups including 10, 15 and 20 authors obtained from a new dataset of newspaper articles. Term frequency-inverse document frequency (TF-IDF) vectors were created by using 1-gram and 2-gram word tokens. Our results show that the most successful method is the SGD with a classification performance accuracy of 0.976% by using word unigrams and most successful method is the LR with a classification performance accuracy of 0.935% by using word bigrams.","PeriodicalId":166049,"journal":{"name":"International Journal of Multidisciplinary Studies and Innovative Technologies","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1900-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Author Identification with Machine Learning Algorithms\",\"authors\":\"İbrahim Yülüce, Feriştah Dalkılıç\",\"doi\":\"10.36287/ijmsit.6.1.45\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"– Author identification is one of the application areas of text mining. It deals with the automatic prediction of the potential author of an electronic text among predefined author candidates by using author specific writing styles. In this study, we conducted an experiment for the identification of the author of a Turkish language text by using classical machine learning methods including Support Vector Machines (SVM), Gaussian Naive Bayes (GaussianNB), Multi Layer Perceptron (MLP), Logistic Regression (LR), Stochastic Gradient Descent (SGD) and ensemble learning methods including Extremely Randomized Trees (ExtraTrees), and eXtreme Gradient Boosting (XGBoost). The proposed method was applied on three different sizes of author groups including 10, 15 and 20 authors obtained from a new dataset of newspaper articles. Term frequency-inverse document frequency (TF-IDF) vectors were created by using 1-gram and 2-gram word tokens. Our results show that the most successful method is the SGD with a classification performance accuracy of 0.976% by using word unigrams and most successful method is the LR with a classification performance accuracy of 0.935% by using word bigrams.\",\"PeriodicalId\":166049,\"journal\":{\"name\":\"International Journal of Multidisciplinary Studies and Innovative Technologies\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"1900-01-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"International Journal of Multidisciplinary Studies and Innovative Technologies\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.36287/ijmsit.6.1.45\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"International Journal of Multidisciplinary Studies and Innovative Technologies","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.36287/ijmsit.6.1.45","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

作者识别是文本挖掘的应用领域之一。它通过使用作者特定的写作风格，在预定义的作者候选者中自动预测电子文本的潜在作者。在这项研究中，我们使用经典的机器学习方法，包括支持向量机(SVM)、高斯朴素贝叶斯(GaussianNB)、多层感知器(MLP)、逻辑回归(LR)、随机梯度下降(SGD)和集成学习方法，包括极端随机树(ExtraTrees)和极端梯度提升(XGBoost)，进行了土耳其语文本作者识别的实验。将该方法应用于从新的报纸文章数据集中获得的三种不同规模的作者组，包括10、15和20名作者。术语频率逆文档频率(TF-IDF)向量是通过使用1克和2克单词标记创建的。我们的研究结果表明，最成功的方法是使用词单图的SGD，其分类性能准确率为0.976%;最成功的方法是使用词双图的LR，其分类性能准确率为0.935%。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Author Identification with Machine Learning Algorithms

– Author identification is one of the application areas of text mining. It deals with the automatic prediction of the potential author of an electronic text among predefined author candidates by using author specific writing styles. In this study, we conducted an experiment for the identification of the author of a Turkish language text by using classical machine learning methods including Support Vector Machines (SVM), Gaussian Naive Bayes (GaussianNB), Multi Layer Perceptron (MLP), Logistic Regression (LR), Stochastic Gradient Descent (SGD) and ensemble learning methods including Extremely Randomized Trees (ExtraTrees), and eXtreme Gradient Boosting (XGBoost). The proposed method was applied on three different sizes of author groups including 10, 15 and 20 authors obtained from a new dataset of newspaper articles. Term frequency-inverse document frequency (TF-IDF) vectors were created by using 1-gram and 2-gram word tokens. Our results show that the most successful method is the SGD with a classification performance accuracy of 0.976% by using word unigrams and most successful method is the LR with a classification performance accuracy of 0.935% by using word bigrams.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

International Journal of Multidisciplinary Studies and Innovative Technologies

自引率

0.00%

发文量