Building a Large Chinese Corpus Annotated with Semantic Dependency

Workshop on Chinese Language Processing Pub Date : 2003-07-11 DOI:10.3115/1119250.1119262

Mingqin Li, Juan-Zi Li, Zhendong Dong, Zuoying Wang, Dajin Lu

引用次数: 30

Abstract

At present most of corpora are annotated mainly with syntactic knowledge. In this paper, we attempt to build a large corpus and annotate semantic knowledge with dependency grammar. We believe that words are the basic units of semantics, and the structure and meaning of a sentence consist mainly of a series of semantic dependencies between individual words. A 1,000,000-word-scale corpus annotated with semantic dependency has been built. Compared with syntactic knowledge, semantic knowledge is more difficult to annotate, for ambiguity problem is more serious. In the paper, the strategy to improve consistency is addressed, and congruence is defined to measure the consistency of tagged corpus.. Finally, we will compare our corpus with other well-known corpora.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

基于语义依赖标注的大型汉语语料库的构建

目前大多数语料库的标注都是以句法知识为主。在本文中，我们尝试建立一个大型语料库，并用依存语法对语义知识进行标注。我们认为词是语义的基本单位，句子的结构和意义主要由单个词之间的一系列语义依赖组成。建立了一个带有语义依赖注释的100万字规模的语料库。与句法知识相比，语义知识的标注难度更大，歧义问题更严重。本文讨论了提高一致性的策略，并定义了一致性来衡量标记语料库的一致性。最后，我们将我们的语料库与其他知名语料库进行比较。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Workshop on Chinese Language Processing

自引率

0.00%

发文量

期刊最新文献

Building a Large Chinese Corpus Annotated with Semantic Dependency A Two-stage Statistical Word Segmentation System for Chinese Unsupervised Training for Overlapping Ambiguity Resolution in Chinese Word Segmentation Chinese Word Segmentation in MSR-NLP Annotating the Propositions in the Penn Chinese Treebank