基于Tesseract OCR和后处理的键值对搜索系统

Áron Zoltán Kaló, M. Sipos
{"title":"基于Tesseract OCR和后处理的键值对搜索系统","authors":"Áron Zoltán Kaló, M. Sipos","doi":"10.1109/SAMI50585.2021.9378680","DOIUrl":null,"url":null,"abstract":"Optical character recognition systems make it possible to extract text from images. In many cases, this may be sufficient, but there are cases where key-value pairs are required. In this paper, we investigate the use of the open source Tesseract OCR system, to extract text data from images, and perform a key-value pair search. Image noise needs to be minimized with image processing algorithms before recognition. It is necessary to perform so-called post processing procedures on the output of the Tesseract. These post-processors can transform the result of the recognition performed by the OCR system. Those can improve the accuracy of the information extracted during the transformation, for example with the help of regular expressions. The key value pair search is performed after these procedures.","PeriodicalId":402414,"journal":{"name":"2021 IEEE 19th World Symposium on Applied Machine Intelligence and Informatics (SAMI)","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2021-01-21","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"3","resultStr":"{\"title\":\"Key-Value Pair Searhing System via Tesseract OCR and Post Processing\",\"authors\":\"Áron Zoltán Kaló, M. Sipos\",\"doi\":\"10.1109/SAMI50585.2021.9378680\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Optical character recognition systems make it possible to extract text from images. In many cases, this may be sufficient, but there are cases where key-value pairs are required. In this paper, we investigate the use of the open source Tesseract OCR system, to extract text data from images, and perform a key-value pair search. Image noise needs to be minimized with image processing algorithms before recognition. It is necessary to perform so-called post processing procedures on the output of the Tesseract. These post-processors can transform the result of the recognition performed by the OCR system. Those can improve the accuracy of the information extracted during the transformation, for example with the help of regular expressions. The key value pair search is performed after these procedures.\",\"PeriodicalId\":402414,\"journal\":{\"name\":\"2021 IEEE 19th World Symposium on Applied Machine Intelligence and Informatics (SAMI)\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2021-01-21\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"3\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2021 IEEE 19th World Symposium on Applied Machine Intelligence and Informatics (SAMI)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/SAMI50585.2021.9378680\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2021 IEEE 19th World Symposium on Applied Machine Intelligence and Informatics (SAMI)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/SAMI50585.2021.9378680","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 3

摘要

光学字符识别系统使从图像中提取文本成为可能。在许多情况下,这可能就足够了,但是在某些情况下需要键值对。在本文中,我们研究了使用开源的Tesseract OCR系统,从图像中提取文本数据,并执行键值对搜索。在识别之前,需要使用图像处理算法将图像噪声降至最低。有必要对Tesseract的输出执行所谓的后处理程序。这些后置处理器可以对OCR系统执行的识别结果进行变换。它们可以提高在转换过程中提取信息的准确性,例如在正则表达式的帮助下。在这些过程之后执行键值对搜索。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
Key-Value Pair Searhing System via Tesseract OCR and Post Processing
Optical character recognition systems make it possible to extract text from images. In many cases, this may be sufficient, but there are cases where key-value pairs are required. In this paper, we investigate the use of the open source Tesseract OCR system, to extract text data from images, and perform a key-value pair search. Image noise needs to be minimized with image processing algorithms before recognition. It is necessary to perform so-called post processing procedures on the output of the Tesseract. These post-processors can transform the result of the recognition performed by the OCR system. Those can improve the accuracy of the information extracted during the transformation, for example with the help of regular expressions. The key value pair search is performed after these procedures.
求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
自引率
0.00%
发文量
0
期刊最新文献
Usage of RAPTOR for travel time minimizing journey planner Slip Control by Identifying the Magnetic Field of the Elements of an Asynchronous Motor Supervised Operational Change Point Detection using Ensemble Long-Short Term Memory in a Multicomponent Industrial System Improving the activity recognition using GMAF and transfer learning in post-stroke rehabilitation assessment A Baseline Assessment Method of UAV Swarm Resilience Based on Complex Networks*
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1