OCR for Greek polytonic (multi accent) historical printed documents: development, optimization and quality control

Proceedings of the 3rd International Conference on Digital Access to Textual Cultural Heritage Pub Date : 2019-05-08 DOI:10.1145/3322905.3322926

Anna-Maria Sichani, Panagiotis Kaddas, Georgios K. Mikros, B. Gatos

引用次数: 1

Abstract

This paper presents the development and implementation of a robust OCR tool and a related comprehensive workflow for the recognition of Greek printed polytonic scripts. This project is initiated and developed by an interdisciplinary team with expertise in the areas of document image processing, character segmentation and recognition, machine learning, corpus creation and digital humanities. Our paper aims to describe the design and development of the workflow around this project, including data gathering and structuring, OCR tool development, user interface development, experiments on the training procedure of the tool, evaluation, post-correction and quality control of the results.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

希腊多音(多口音)历史印刷文献的OCR:开发、优化和质量控制

本文介绍了开发和实现一个强大的OCR工具和相关的综合工作流程，用于识别希腊印刷多音脚本。该项目由一个跨学科团队发起和开发，该团队在文档图像处理、字符分割和识别、机器学习、语料库创建和数字人文等领域具有专业知识。本文旨在描述围绕该项目的工作流程的设计和开发，包括数据收集和构建，OCR工具开发，用户界面开发，工具训练程序实验，结果的评估，后校正和质量控制。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Proceedings of the 3rd International Conference on Digital Access to Textual Cultural Heritage

自引率

0.00%

发文量