If You Build Your Own NER Scorer, Non-replicable Results Will Come

First Workshop on Insights from Negative Results in NLP Pub Date : 2020-11-01 DOI:10.18653/v1/2020.insights-1.15

Constantine Lignos, Marjan Kamyab

引用次数: 7

Abstract

We attempt to replicate a named entity recognition (NER) model implemented in a popular toolkit and discover that a critical barrier to doing so is the inconsistent evaluation of improper label sequences. We define these sequences and examine how two scorers differ in their handling of them, finding that one approach produces F1 scores approximately 0.5 points higher on the CoNLL 2003 English development and test sets. We propose best practices to increase the replicability of NER evaluations by increasing transparency regarding the handling of improper label sequences.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

如果你建立自己的NER记分器，不可复制的结果将会到来

我们试图复制一个在流行工具包中实现的命名实体识别(NER)模型，并发现这样做的一个关键障碍是对不适当的标签序列的不一致评估。我们定义了这些序列，并检查了两个评分者在处理它们时的差异，发现一种方法在CoNLL 2003英语发展和测试集上产生的F1分数大约高出0.5分。我们提出了最佳实践，通过增加处理不当标签序列的透明度来提高NER评估的可复制性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

First Workshop on Insights from Negative Results in NLP

自引率

0.00%

发文量

期刊最新文献

What GPT Knows About Who is Who Pathologies of Pre-trained Language Models in Few-shot Fine-tuning Can Question Rewriting Help Conversational Question Answering? Extending the Scope of Out-of-Domain: Examining QA models in multiple subdomains Do Data-based Curricula Work?