Abstract argumentation for reading order detection

Proceedings of the ACM Symposium on Document Engineering. ACM Symposium on Document Engineering Pub Date : 2014-09-16 DOI:10.1145/2644866.2644883

S. Ferilli, D. Grieco, Domenico Redavid, F. Esposito

引用次数: 4

Abstract

Detecting the reading order among the layout components of a document's page is fundamental to ensure effectiveness or even applicability of subsequent content extraction steps. While in single-column documents the reading flow can be straightforwardly determined, in more complex documents the task may become very hard. This paper proposes an automatic strategy for identifying the correct reading order of a document page's components based on abstract argumentation. The technique is unsupervised, and works on any kind of document based only on general assumptions about how humans behave when reading documents. Experimental results show that it is effective in more complex cases, and requires less background knowledge, than previous solutions that have been proposed in the literature.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

阅读顺序检测的抽象论证

检测文档页面布局组件之间的阅读顺序是确保后续内容提取步骤的有效性甚至适用性的基础。在单列文档中，可以直接确定阅读流程，但在更复杂的文档中，这项任务可能会变得非常困难。本文提出了一种基于抽象论证的文档页面组件正确阅读顺序自动识别策略。该技术是无监督的，并且仅基于人类在阅读文档时的行为的一般假设来处理任何类型的文档。实验结果表明，与文献中提出的解决方案相比，该方法在更复杂的情况下是有效的，并且需要更少的背景知识。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Proceedings of the ACM Symposium on Document Engineering. ACM Symposium on Document Engineering

自引率

0.00%

发文量

期刊最新文献

The Notarial Archives, Valletta: Starting from Zero Truncation: all the news that fits we'll print Classifying and ranking search engine results as potential sources of plagiarism An ensemble approach for text document clustering using Wikipedia concepts Document changes: modeling, detection, storage and visualization (DChanges 2014)