Word searching in CCITT group 4 compressed document images

Seventh International Conference on Document Analysis and Recognition, 2003. Proceedings. Pub Date : 2003-08-03 DOI:10.1109/ICDAR.2003.1227709

Yue Lu, C. Tan

引用次数: 16

Abstract

In this paper, we present a compressed pattern matching method for searching user queried words in the CCITT Group 4 compressed document images, without decompressing. The feature pixels composed of black changing elements and white changing elements are extracted directly from the CCITT Group 4 compressed document images. The connected components are labeled based on a line-by-line strategy according to the relative positions between the changing elements of the current coding line and the changing elements of the reference line. Word boxes are bounded by merging the connected components. A two-stage matching strategy is constructed to measure the dissimilarity between the template image of the user's query word and the words extracted from document images. Experimental results confirmed the validity of the proposed approach.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

CCITT组4压缩文档图像中的单词搜索

在本文中，我们提出了一种压缩模式匹配方法，用于在CCITT Group 4压缩文档图像中搜索用户查询的单词，而无需解压。直接从CCITT Group 4压缩文档图像中提取由黑色变化元素和白色变化元素组成的特征像素。根据当前编码线的变化元素与参考线的变化元素之间的相对位置，基于逐行策略对所连接的组件进行标记。单词框通过合并连接的组件来限定。构造了一种两阶段匹配策略来度量用户查询词的模板图像与从文档图像中提取的词之间的不相似度。实验结果证实了该方法的有效性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Seventh International Conference on Document Analysis and Recognition, 2003. Proceedings.

自引率

0.00%

发文量

期刊最新文献

Impact of imperfect OCR on part-of-speech tagging Writer identification using innovative binarised features of handwritten numerals Word searching in CCITT group 4 compressed document images Exploiting reliability for dynamic selection of classi .ers by means of genetic algorithms Investigation of off-line Japanese signature verification using a pattern matching