Uncovering Surprising Event Boundaries in Narratives

Proceedings of the 4th Workshop of Narrative Understanding (WNU2022) Pub Date : 1900-01-01 DOI:10.18653/v1/2022.wnu-1.1

Zhiling Wang, A. Jafarpour, Maarten Sap

引用次数: 4

Abstract

It is important to define meaningful and interpretable automatic evaluation metrics for open-domain dialog research. Standard language generation metrics have been shown to be ineffective for dialog. This paper introduces the FED metric (fine-grained evaluation of dialog), an automatic evaluation metric which uses DialoGPT, without any fine-tuning or supervision. It also introduces the FED dataset which is constructed by annotating a set of human-system and human-human conversations with eighteen fine-grained dialog qualities. The FED metric (1) does not rely on a ground-truth response, (2) does not require training data and (3) measures fine-grained dialog qualities at both the turn and whole dialog levels. FED attains moderate to strong correlation with human judgement at both levels.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

揭示叙事中令人惊讶的事件边界

在开放域对话研究中，定义有意义且可解释的自动评价指标是非常重要的。标准语言生成度量已被证明对对话是无效的。本文介绍了FED度量(细粒度的对话评估)，这是一种使用DialoGPT的自动评估度量，不需要任何微调和监督。它还介绍了FED数据集，该数据集通过注释一组具有18个细粒度对话质量的人-系统和人-人对话来构建。FED度量(1)不依赖于真实的响应，(2)不需要训练数据，(3)在回合和整个对话级别测量细粒度的对话质量。FED与人的判断在这两个层面上都达到了中度到高度的相关性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Proceedings of the 4th Workshop of Narrative Understanding (WNU2022)

自引率

0.00%

发文量

期刊最新文献

GPT-2-based Human-in-the-loop Theatre Play Script Generation Narrative Detection and Feature Analysis in Online Health Communities Compositional Generalization for Kinship Prediction through Data Augmentation Looking from the Inside: How Children Render Character’s Perspectives in Freely Told Fantasy Stories How to be Helpful on Online Support Forums?