FantasyCoref: Coreference Resolution on Fantasy Literature Through Omniscient Writer’s Point of View

Sooyoun Han, Sumin Seo, Mi-gyu Kang, Jongin Kim, Nayoung Choi, Minyue Song, Jinho D. Choi
{"title":"FantasyCoref: Coreference Resolution on Fantasy Literature Through Omniscient Writer’s Point of View","authors":"Sooyoun Han, Sumin Seo, Mi-gyu Kang, Jongin Kim, Nayoung Choi, Minyue Song, Jinho D. Choi","doi":"10.18653/v1/2021.crac-1.3","DOIUrl":null,"url":null,"abstract":"This paper presents a new corpus and annotation guideline for a novel coreference resolution task on fictional texts, and analyzes its unique characteristics. FantasyCoref contains 211 stories of Grimms’ Fairy Tales and 3 other fantasy literature annotated in the omniscient writer’s point of view (OWV) to handle distinctive aspects in this genre. This task is more challenging than general coreference resolution in two ways. First, documents in our corpus are 2.5 times longer than the ones in OntoNotes, raising a new layer of difficulty in resolving long-distant referents. Second, annotation of literary styles and concepts raise several issues which are not sufficiently addressed in the existing annotation guidelines. Hence, considerations on such issues and the concept of OWV are necessary to achieve high inter-annotator agreement (IAA) in coreference resolution of fictional texts. We carefully conduct annotation tasks in four stages to ensure the quality of our annotation. As a result, a high IAA score of 87% is achieved using the standard coreference evaluation metric. Finally, state-of-the-art coreference resolution approaches are evaluated on our corpus. After training with our annotated dataset, there was a 2.59% and 3.06% improvement over the model trained on the OntoNotes dataset. Also, we observe that the portion of errors specific to fictional texts declines after the training.","PeriodicalId":447425,"journal":{"name":"Proceedings of the Fourth Workshop on Computational Models of Reference, Anaphora and Coreference","volume":"11 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1900-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the Fourth Workshop on Computational Models of Reference, Anaphora and Coreference","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.18653/v1/2021.crac-1.3","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 2

Abstract

This paper presents a new corpus and annotation guideline for a novel coreference resolution task on fictional texts, and analyzes its unique characteristics. FantasyCoref contains 211 stories of Grimms’ Fairy Tales and 3 other fantasy literature annotated in the omniscient writer’s point of view (OWV) to handle distinctive aspects in this genre. This task is more challenging than general coreference resolution in two ways. First, documents in our corpus are 2.5 times longer than the ones in OntoNotes, raising a new layer of difficulty in resolving long-distant referents. Second, annotation of literary styles and concepts raise several issues which are not sufficiently addressed in the existing annotation guidelines. Hence, considerations on such issues and the concept of OWV are necessary to achieve high inter-annotator agreement (IAA) in coreference resolution of fictional texts. We carefully conduct annotation tasks in four stages to ensure the quality of our annotation. As a result, a high IAA score of 87% is achieved using the standard coreference evaluation metric. Finally, state-of-the-art coreference resolution approaches are evaluated on our corpus. After training with our annotated dataset, there was a 2.59% and 3.06% improvement over the model trained on the OntoNotes dataset. Also, we observe that the portion of errors specific to fictional texts declines after the training.
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
从无所不知的作家的角度看奇幻文学的共涉解析
针对一种新型的虚构文本共指解析任务,提出了一种新的语料库和标注准则,并分析了其独特的特点。FantasyCoref收录了格林童话的211个故事和其他3个奇幻文学作品,以全知作家的观点(OWV)来处理这一类型的独特方面。这一任务比一般的共参解析在两个方面更具挑战性。首先,我们语料库中的文档比OntoNotes中的文档长2.5倍,这增加了解析远距离引用的难度。其次,文学风格和概念的注释提出了一些问题,这些问题在现有的注释指南中没有得到充分的解决。因此,为了在虚构文本的共指解析中实现高度的注释者间一致性(IAA),有必要考虑这些问题和OWV的概念。我们分四个阶段认真进行标注任务,确保标注质量。结果,使用标准的共参考评价指标,获得了87%的高IAA评分。最后,在我们的语料库上评估了最先进的共参考解析方法。在使用我们的注释数据集进行训练后,与在OntoNotes数据集上训练的模型相比,分别有2.59%和3.06%的改进。此外,我们观察到,在训练后,特定于虚构文本的错误比例下降。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
自引率
0.00%
发文量
0
期刊最新文献
Understanding Mention Detector-Linker Interaction in Neural Coreference Resolution DramaCoref: A Hybrid Coreference Resolution System for German Theater Plays Resources and Evaluations for Danish Entity Resolution Event and Entity Coreference using Trees to Encode Uncertainty in Joint Decisions FantasyCoref: Coreference Resolution on Fantasy Literature Through Omniscient Writer’s Point of View
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1