Table Recognition and Understanding from PDF Files

Ninth International Conference on Document Analysis and Recognition (ICDAR 2007) Pub Date : 2007-09-23 DOI:10.1109/ICDAR.2007.241

Tamir Hassan, Robert Baumgartner

引用次数: 70

Abstract

We propose a flexible method for detecting and understanding tables in PDF files, which is not reliant upon one particular feature being present, for example ruling lines or indentations, and is therefore applicable to a wide variety of visual presentations. We describe the steps required in transforming the low-level PDF instructions into text segments, lines and boxes on a page. We propose three different classifications for published tables, and develop methods to detect these tables and correctly identify their respective rows and columns. We also explain how to recognize spanning rows and columns, and multi-line rows. Experimental results show that our algorithm is effective in converting a wide variety of tabular presentations into HTML for information extraction purposes.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

从PDF文件中识别和理解表

我们提出了一种灵活的方法来检测和理解PDF文件中的表，这种方法不依赖于存在的特定特性，例如规则行或缩进，因此适用于各种视觉表示。我们描述了将低级PDF指令转换为页面上的文本段、行和框所需的步骤。我们对已发布的表提出了三种不同的分类，并开发了检测这些表并正确识别其各自行和列的方法。我们还解释了如何识别跨行、跨列以及多行。实验结果表明，该算法可以有效地将各种表格表示转换为HTML以用于信息提取。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Ninth International Conference on Document Analysis and Recognition (ICDAR 2007)

自引率

0.00%

发文量