{"title":"Text area detection in handwritten documents scanned for further processing","authors":"J. Pach, Izabella Antoniuk, A. Krupa","doi":"10.22630/mgv.2020.29.1.2","DOIUrl":null,"url":null,"abstract":"In this paper we present an approach to text area detection using binary images, Constrained Run Length Algorithm and other noise reduction methods of removing the artefacts. Text processing includes various activities, most of which are related to preparing input data for further operations in the best possible way, that will not hinder the OCR algorithms. This is especially the case when handwritten manuscripts are considered, and even more so with very old documents. We present our methodology for text area detection problem, which is capable of removing most of irrelevant objects, including elements such as page edges, stains, folds etc. At the same time the presented method can handle multi-column texts or varying line thickness. The generated mask can accurately mark the actual text area, so that the output image can be easily used in further text processing steps.","PeriodicalId":39750,"journal":{"name":"Machine Graphics and Vision","volume":"55 1","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2020-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Machine Graphics and Vision","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.22630/mgv.2020.29.1.2","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0
Abstract
In this paper we present an approach to text area detection using binary images, Constrained Run Length Algorithm and other noise reduction methods of removing the artefacts. Text processing includes various activities, most of which are related to preparing input data for further operations in the best possible way, that will not hinder the OCR algorithms. This is especially the case when handwritten manuscripts are considered, and even more so with very old documents. We present our methodology for text area detection problem, which is capable of removing most of irrelevant objects, including elements such as page edges, stains, folds etc. At the same time the presented method can handle multi-column texts or varying line thickness. The generated mask can accurately mark the actual text area, so that the output image can be easily used in further text processing steps.
期刊介绍:
Machine GRAPHICS & VISION (MGV) is a refereed international journal, published quarterly, providing a scientific exchange forum and an authoritative source of information in the field of, in general, pictorial information exchange between computers and their environment, including applications of visual and graphical computer systems. The journal concentrates on theoretical and computational models underlying computer generated, analysed, or otherwise processed imagery, in particular: - image processing - scene analysis, modeling, and understanding - machine vision - pattern matching and pattern recognition - image synthesis, including three-dimensional imaging and solid modeling