Nowhere to Hide: Finding Plagiarized Documents Based on Sentence Similarity

2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology Pub Date : 2008-12-09 DOI:10.1109/WIIAT.2008.16

Nathaniel Gustafson, M. S. Pera, Yiu-Kai Ng

引用次数: 27

Abstract

Plagiarism is a serious problem that infringes copyrighted documents/materials, which is an unethical practice and decreases the economic incentive received by authors (owners) of the original copies. Unfortunately, plagiarism is getting worse due to the increasing number of on-line publications on the Web, which facilitates locating and paraphrasing information. In solving this problem, we propose a novel plagiarism-detection method, called SimPaD, which (i) establishes the degree of resemblance between any two documents D1 and D2 based on their sentence-to-sentence similarity computed by using pre-defined word-correlation factors, and (ii) generates agraphical view of sentences that are similar (or the same) in D1 and D2. Experimental results verify that SimPaD is highly accurate in detecting (non-) plagiarized documents and outperforms existing plagiarism-detection approaches.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

无处可藏:基于句子相似度查找剽窃文档

抄袭是一个严重的问题，它侵犯了受版权保护的文件/材料，这是一种不道德的做法，并减少了原始副本的作者(所有者)获得的经济激励。不幸的是，由于网络上的在线出版物越来越多，这使得定位和解释信息变得更加容易，剽窃现象越来越严重。为了解决这一问题，我们提出了一种名为SimPaD的新型剽窃检测方法，该方法(i)通过使用预定义的单词相关因子计算任意两个文档D1和D2之间的句子相似度，并(ii)生成D1和D2中相似(或相同)句子的图形视图。实验结果表明SimPaD在检测(非)剽窃文档方面具有很高的准确性，并且优于现有的剽窃检测方法。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2008 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology

自引率

0.00%

发文量

期刊最新文献

Effective Usage of Computational Trust Models in Rational Environments Link-Based Anomaly Detection in Communication Networks Quality Information Retrieval for the World Wide Web A k-Nearest-Neighbour Method for Classifying Web Search Results with Data in Folksonomies Concept Extraction and Clustering for Topic Digital Library Construction