无处可藏:基于句子相似度查找剽窃文档

2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology Pub Date : 2008-12-09 DOI:10.1109/WIIAT.2008.16

Nathaniel Gustafson, M. S. Pera, Yiu-Kai Ng

{"title":"无处可藏:基于句子相似度查找剽窃文档","authors":"Nathaniel Gustafson, M. S. Pera, Yiu-Kai Ng","doi":"10.1109/WIIAT.2008.16","DOIUrl":null,"url":null,"abstract":"Plagiarism is a serious problem that infringes copyrighted documents/materials, which is an unethical practice and decreases the economic incentive received by authors (owners) of the original copies. Unfortunately, plagiarism is getting worse due to the increasing number of on-line publications on the Web, which facilitates locating and paraphrasing information. In solving this problem, we propose a novel plagiarism-detection method, called SimPaD, which (i) establishes the degree of resemblance between any two documents D1 and D2 based on their sentence-to-sentence similarity computed by using pre-defined word-correlation factors, and (ii) generates agraphical view of sentences that are similar (or the same) in D1 and D2. Experimental results verify that SimPaD is highly accurate in detecting (non-) plagiarized documents and outperforms existing plagiarism-detection approaches.","PeriodicalId":393772,"journal":{"name":"2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology","volume":null,"pages":null},"PeriodicalIF":0.0000,"publicationDate":"2008-12-09","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"27","resultStr":"{\"title\":\"Nowhere to Hide: Finding Plagiarized Documents Based on Sentence Similarity\",\"authors\":\"Nathaniel Gustafson, M. S. Pera, Yiu-Kai Ng\",\"doi\":\"10.1109/WIIAT.2008.16\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Plagiarism is a serious problem that infringes copyrighted documents/materials, which is an unethical practice and decreases the economic incentive received by authors (owners) of the original copies. Unfortunately, plagiarism is getting worse due to the increasing number of on-line publications on the Web, which facilitates locating and paraphrasing information. In solving this problem, we propose a novel plagiarism-detection method, called SimPaD, which (i) establishes the degree of resemblance between any two documents D1 and D2 based on their sentence-to-sentence similarity computed by using pre-defined word-correlation factors, and (ii) generates agraphical view of sentences that are similar (or the same) in D1 and D2. Experimental results verify that SimPaD is highly accurate in detecting (non-) plagiarized documents and outperforms existing plagiarism-detection approaches.\",\"PeriodicalId\":393772,\"journal\":{\"name\":\"2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology\",\"volume\":null,\"pages\":null},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2008-12-09\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"27\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/WIIAT.2008.16\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/WIIAT.2008.16","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 27

摘要

抄袭是一个严重的问题，它侵犯了受版权保护的文件/材料，这是一种不道德的做法，并减少了原始副本的作者(所有者)获得的经济激励。不幸的是，由于网络上的在线出版物越来越多，这使得定位和解释信息变得更加容易，剽窃现象越来越严重。为了解决这一问题，我们提出了一种名为SimPaD的新型剽窃检测方法，该方法(i)通过使用预定义的单词相关因子计算任意两个文档D1和D2之间的句子相似度，并(ii)生成D1和D2中相似(或相同)句子的图形视图。实验结果表明SimPaD在检测(非)剽窃文档方面具有很高的准确性，并且优于现有的剽窃检测方法。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Nowhere to Hide: Finding Plagiarized Documents Based on Sentence Similarity

Plagiarism is a serious problem that infringes copyrighted documents/materials, which is an unethical practice and decreases the economic incentive received by authors (owners) of the original copies. Unfortunately, plagiarism is getting worse due to the increasing number of on-line publications on the Web, which facilitates locating and paraphrasing information. In solving this problem, we propose a novel plagiarism-detection method, called SimPaD, which (i) establishes the degree of resemblance between any two documents D1 and D2 based on their sentence-to-sentence similarity computed by using pre-defined word-correlation factors, and (ii) generates agraphical view of sentences that are similar (or the same) in D1 and D2. Experimental results verify that SimPaD is highly accurate in detecting (non-) plagiarized documents and outperforms existing plagiarism-detection approaches.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology

自引率

0.00%

发文量

期刊最新文献

Effective Usage of Computational Trust Models in Rational Environments Link-Based Anomaly Detection in Communication Networks Quality Information Retrieval for the World Wide Web A k-Nearest-Neighbour Method for Classifying Web Search Results with Data in Folksonomies Concept Extraction and Clustering for Topic Digital Library Construction