{"title":"Document Copy Detection Using the Improved Fuzzy Hashing","authors":"Guohua Wu, Ershuai Fu, Liuyang Wang, Mengmeng Zhao","doi":"10.1109/CSMA.2015.18","DOIUrl":null,"url":null,"abstract":"Document copy detection is an effective method that can protect intellectual property rights as well as improve the efficiency of information retrieval. To our knowledge, it is a common method that using the fingerprints of one document in the process of detecting. Therefore, selecting the appropriate document fingerprints plays a key role. This paper firstly describes several mature methods of selecting document fingerprints, and analyzes their merit and demerit. Then we review the principle of Fuzzy Hashing, which suffers from the instability and inefficiency of fragmenting. To resolve the critical problems, we finally propose a novel algorithm based on the Fuzzy Hashing. Compared to original method, the proposed document copy detection algorithm can not only ensure the proper size of fragment but also enhance the speed of fragmenting. And in terms of efficiency and accuracy, the algorithm achieves high performance.","PeriodicalId":205396,"journal":{"name":"2015 International Conference on Computer Science and Mechanical Automation (CSMA)","volume":"55 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2015-10-23","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2015 International Conference on Computer Science and Mechanical Automation (CSMA)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/CSMA.2015.18","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0
Abstract
Document copy detection is an effective method that can protect intellectual property rights as well as improve the efficiency of information retrieval. To our knowledge, it is a common method that using the fingerprints of one document in the process of detecting. Therefore, selecting the appropriate document fingerprints plays a key role. This paper firstly describes several mature methods of selecting document fingerprints, and analyzes their merit and demerit. Then we review the principle of Fuzzy Hashing, which suffers from the instability and inefficiency of fragmenting. To resolve the critical problems, we finally propose a novel algorithm based on the Fuzzy Hashing. Compared to original method, the proposed document copy detection algorithm can not only ensure the proper size of fragment but also enhance the speed of fragmenting. And in terms of efficiency and accuracy, the algorithm achieves high performance.