{"title":"Localization, extraction and recognition of text in Telugu document images","authors":"A. Negi, K. Shanker, C. K. Chereddi","doi":"10.1109/ICDAR.2003.1227846","DOIUrl":null,"url":null,"abstract":"In this paper we present a system to locate, extract andrecognize Telugu text. The circular nature of Telugu scriptis exploited for segmenting text regions using the HoughTransform. First, the Hough Transform for circles is performedon the Sobel gradient magnitude of the image tolocate text. The located circles are filled to yield text regions,followed by Recursive XY Cuts to segment the regionsinto paragraphs, lines and word regions. A regionmerging process with a bottom-up approach envelopes individualwords. Local binarization of the word MBRs yieldsconnected components containing glyphs for recognition.The recognition process first identifies candidate charactersby a zoning technique and then constructs structural featurevectors by cavity analysis. Finally, if required, crossingcount based non-linear normalization and scaling is performedbefore template matching. The segmentation processsucceeds in extracting text from images with complexNon-Manhattan layouts. The recognition process gave acharacter recognition accuracy of 97%-98%.","PeriodicalId":249193,"journal":{"name":"Seventh International Conference on Document Analysis and Recognition, 2003. Proceedings.","volume":"37 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2003-08-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"31","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Seventh International Conference on Document Analysis and Recognition, 2003. Proceedings.","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICDAR.2003.1227846","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 31
Abstract
In this paper we present a system to locate, extract andrecognize Telugu text. The circular nature of Telugu scriptis exploited for segmenting text regions using the HoughTransform. First, the Hough Transform for circles is performedon the Sobel gradient magnitude of the image tolocate text. The located circles are filled to yield text regions,followed by Recursive XY Cuts to segment the regionsinto paragraphs, lines and word regions. A regionmerging process with a bottom-up approach envelopes individualwords. Local binarization of the word MBRs yieldsconnected components containing glyphs for recognition.The recognition process first identifies candidate charactersby a zoning technique and then constructs structural featurevectors by cavity analysis. Finally, if required, crossingcount based non-linear normalization and scaling is performedbefore template matching. The segmentation processsucceeds in extracting text from images with complexNon-Manhattan layouts. The recognition process gave acharacter recognition accuracy of 97%-98%.