余弦相似度在注释自动分类器中的实现

JISKA Jurnal Informatika Sunan Kalijaga Pub Date : 2019-06-11 DOI:10.14421/JISKA.2018.32-05

Muhammad Habibi

{"title":"余弦相似度在注释自动分类器中的实现","authors":"Muhammad Habibi","doi":"10.14421/JISKA.2018.32-05","DOIUrl":null,"url":null,"abstract":"Classification of text with a large amount is needed to extract the information contained in it. Student comments containing suggestions and criticisms about the lecturer and the lecture process on the learning evaluation system are not well classified, resulting in a difficult assessment process. So from that, we need a classification model that can classify comments automatically into classification categories. The method used is the Cosine Similarity method, which is a method for calculating similarities between two objects expressed in two vectors. The data used in this study were 1,630 comment data with several different categories. The test in this study uses k-fold cross-validation with k = 10. The results showed that the percentage accuracy of the classification model was 80.87%.","PeriodicalId":34216,"journal":{"name":"JISKA Jurnal Informatika Sunan Kalijaga","volume":" ","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2019-06-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"8","resultStr":"{\"title\":\"Implementation of Cosine Similarity in an automatic classifier for comments\",\"authors\":\"Muhammad Habibi\",\"doi\":\"10.14421/JISKA.2018.32-05\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Classification of text with a large amount is needed to extract the information contained in it. Student comments containing suggestions and criticisms about the lecturer and the lecture process on the learning evaluation system are not well classified, resulting in a difficult assessment process. So from that, we need a classification model that can classify comments automatically into classification categories. The method used is the Cosine Similarity method, which is a method for calculating similarities between two objects expressed in two vectors. The data used in this study were 1,630 comment data with several different categories. The test in this study uses k-fold cross-validation with k = 10. The results showed that the percentage accuracy of the classification model was 80.87%.\",\"PeriodicalId\":34216,\"journal\":{\"name\":\"JISKA Jurnal Informatika Sunan Kalijaga\",\"volume\":\" \",\"pages\":\"\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2019-06-11\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"8\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"JISKA Jurnal Informatika Sunan Kalijaga\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.14421/JISKA.2018.32-05\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"JISKA Jurnal Informatika Sunan Kalijaga","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.14421/JISKA.2018.32-05","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 8

摘要

需要对量大的文本进行分类，提取其中包含的信息。学生对讲师和授课过程的意见没有很好地分类，导致评估过程困难。因此，我们需要一个分类模型，可以自动将评论分类到分类类别中。使用的方法是余弦相似度法，这是一种计算用两个向量表示的两个对象之间的相似度的方法。本研究使用的数据是1630个不同类别的评论数据。本研究采用k-fold交叉验证，k = 10。结果表明，该分类模型的准确率为80.87%。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Implementation of Cosine Similarity in an automatic classifier for comments

Classification of text with a large amount is needed to extract the information contained in it. Student comments containing suggestions and criticisms about the lecturer and the lecture process on the learning evaluation system are not well classified, resulting in a difficult assessment process. So from that, we need a classification model that can classify comments automatically into classification categories. The method used is the Cosine Similarity method, which is a method for calculating similarities between two objects expressed in two vectors. The data used in this study were 1,630 comment data with several different categories. The test in this study uses k-fold cross-validation with k = 10. The results showed that the percentage accuracy of the classification model was 80.87%.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

JISKA Jurnal Informatika Sunan Kalijaga

自引率

0.00%

发文量

审稿时长

12 weeks