More versatile scientific documents

Proceedings of the Fourth International Conference on Document Analysis and Recognition Pub Date : 1997-08-18 DOI:10.1109/ICDAR.1997.620680

R. Fateman

{"title":"More versatile scientific documents","authors":"R. Fateman","doi":"10.1109/ICDAR.1997.620680","DOIUrl":null,"url":null,"abstract":"The electronic representation of scientific documents (journals, technical reports, program documentation, laboratory notebooks, etc.) presents challenges in several distinct communities. We see five distinct groups who are concerned with electronic versions of scientific documents: (1) publishers of journals, texts and reference works, and their authors; (2) software publishers for OCR/document analysis and document formatting; (3) software publishers whose products access \"contents semantics\" from documents, including library keyword search programs, natural language search programs, database systems, visual presentation systems, mathematical computation systems, etc.; (4) institutions maintaining access to electronic libraries, which must be broadly construed to include data and programs of all sorts; and (5) individuals and programs acting as their agents who need to use these libraries to identify, locate and retrieve relevant documents. It would be good to have a convergence in design and standards for encoding new or pre-existing (typically paper-based) documents in order to meet the needs of all these groups. Various efforts, some loosely coordinated, but just as often competing, are trying to set standards and build tools. This paper discusses where we are headed.","PeriodicalId":435320,"journal":{"name":"Proceedings of the Fourth International Conference on Document Analysis and Recognition","volume":"30 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1997-08-18","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"9","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the Fourth International Conference on Document Analysis and Recognition","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICDAR.1997.620680","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 9

Abstract

The electronic representation of scientific documents (journals, technical reports, program documentation, laboratory notebooks, etc.) presents challenges in several distinct communities. We see five distinct groups who are concerned with electronic versions of scientific documents: (1) publishers of journals, texts and reference works, and their authors; (2) software publishers for OCR/document analysis and document formatting; (3) software publishers whose products access "contents semantics" from documents, including library keyword search programs, natural language search programs, database systems, visual presentation systems, mathematical computation systems, etc.; (4) institutions maintaining access to electronic libraries, which must be broadly construed to include data and programs of all sorts; and (5) individuals and programs acting as their agents who need to use these libraries to identify, locate and retrieve relevant documents. It would be good to have a convergence in design and standards for encoding new or pre-existing (typically paper-based) documents in order to meet the needs of all these groups. Various efforts, some loosely coordinated, but just as often competing, are trying to set standards and build tools. This paper discusses where we are headed.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

更多样化的科学文献

科学文献(期刊、技术报告、程序文档、实验室笔记等)的电子表示在几个不同的群体中提出了挑战。我们看到有五个不同的群体关注科学文献的电子版本:(1)期刊、文本和参考文献的出版商及其作者;(2)用于OCR/文档分析和文档格式化的软件出版商;(三)其产品从文档中获取“内容语义”的软件发布者，包括图书馆关键字搜索程序、自然语言搜索程序、数据库系统、可视化呈现系统、数学计算系统等;(4)维护电子图书馆访问权限的机构，电子图书馆必须广义地理解为包括各种数据和程序;(5)需要使用这些库来识别、定位和检索相关文档的个人和程序。为了满足所有这些群体的需求，最好在编码新文档或已有文档(通常是基于纸张的)的设计和标准方面有一个统一。各种各样的努力，有些是松散协调的，但也经常是相互竞争的，都在试图设定标准和构建工具。本文讨论了我们的发展方向。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Proceedings of the Fourth International Conference on Document Analysis and Recognition

自引率

0.00%

发文量

期刊最新文献

Document layout analysis based on emergent computation Offline handwritten Chinese character recognition via radical extraction and recognition Boundary normalization for recognition of non-touching non-degraded characters Words recognition using associative memory Image and text coupling for creating electronic books from manuscripts