文献互助智能选刊最新文献

高级搜索发布求助登录注册

Effectively and efficiently detect web page duplication

2009 Fourth International Conference on Digital Information Management Pub Date : 2009-12-18 DOI:10.1109/ICDIM.2009.5356801

Zhongming Han, Qian Mo, Hongzhi Liu, Jianzhi Sun

引用次数: 7

Abstract

There are a lot of redundant web pages on Internet. Based on tag statistic and text similarity comparison, we present a novel multilayer framework for detecting duplicated web pages in this paper. We propose two similarity text paragraphs detection algorithms and implement our framework. The experimental results show that our approach achieves high performance, which means that duplicated web pages can be efficiently detected simply by tag statistic and text comparison.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

有效和高效地检测网页重复

互联网上有很多冗余的网页。本文基于标签统计和文本相似度比较，提出了一种新的多层网页重复检测框架。我们提出了两种相似文本段落检测算法并实现了我们的框架。实验结果表明，该方法取得了较高的性能，仅通过标记统计和文本比较就能有效地检测出重复的网页。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2009 Fourth International Conference on Digital Information Management

2009 Fourth International Conference on Digital Information Management

自引率

0.00%

发文量

0

期刊最新文献

Ontology based entity disambiguation with natural language patterns Tiles — A model for classifying and using contextual information for context-aware applications Effectively and efficiently detect web page duplication From state-based to event-based contextual security policies P2P applied in CMS for advertising

0

微信

客服QQ

Book学术公众号

扫码关注我们

反馈

Book学术官方微信

Book学术文献互助

Book学术文献互助群
群号：604180095

文献互助智能选刊最新文献互助须知联系我们：info@booksci.cn

Book学术提供免费学术资源搜索服务，方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。

Copyright © 2023 Book学术 All rights reserved.

京公网安备 11010802042870号京ICP备2023020795号-1