{"title":"Word searching in CCITT group 4 compressed document images","authors":"Yue Lu, C. Tan","doi":"10.1109/ICDAR.2003.1227709","DOIUrl":null,"url":null,"abstract":"In this paper, we present a compressed pattern matching method for searching user queried words in the CCITT Group 4 compressed document images, without decompressing. The feature pixels composed of black changing elements and white changing elements are extracted directly from the CCITT Group 4 compressed document images. The connected components are labeled based on a line-by-line strategy according to the relative positions between the changing elements of the current coding line and the changing elements of the reference line. Word boxes are bounded by merging the connected components. A two-stage matching strategy is constructed to measure the dissimilarity between the template image of the user's query word and the words extracted from document images. Experimental results confirmed the validity of the proposed approach.","PeriodicalId":249193,"journal":{"name":"Seventh International Conference on Document Analysis and Recognition, 2003. Proceedings.","volume":"58 ","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2003-08-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"16","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Seventh International Conference on Document Analysis and Recognition, 2003. Proceedings.","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICDAR.2003.1227709","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 16
Abstract
In this paper, we present a compressed pattern matching method for searching user queried words in the CCITT Group 4 compressed document images, without decompressing. The feature pixels composed of black changing elements and white changing elements are extracted directly from the CCITT Group 4 compressed document images. The connected components are labeled based on a line-by-line strategy according to the relative positions between the changing elements of the current coding line and the changing elements of the reference line. Word boxes are bounded by merging the connected components. A two-stage matching strategy is constructed to measure the dissimilarity between the template image of the user's query word and the words extracted from document images. Experimental results confirmed the validity of the proposed approach.
在本文中,我们提出了一种压缩模式匹配方法,用于在CCITT Group 4压缩文档图像中搜索用户查询的单词,而无需解压。直接从CCITT Group 4压缩文档图像中提取由黑色变化元素和白色变化元素组成的特征像素。根据当前编码线的变化元素与参考线的变化元素之间的相对位置,基于逐行策略对所连接的组件进行标记。单词框通过合并连接的组件来限定。构造了一种两阶段匹配策略来度量用户查询词的模板图像与从文档图像中提取的词之间的不相似度。实验结果证实了该方法的有效性。