A Perplexity-Based Method for Similar Languages Discrimination

Workshop on NLP for Similar Languages, Varieties and Dialects Pub Date : 1900-01-01 DOI:10.18653/v1/W17-1213

Pablo Gamallo, José Ramom Pichel Campos, I. Alegria

引用次数: 23

Abstract

This article describes the system submitted by the Citius_Ixa_Imaxin team to the VarDial 2017 (DSL and GDI tasks). The strategy underlying our system is based on a language distance computed by means of model perplexity. The best model configuration we have tested is a voting system making use of several n-grams models of both words and characters, even if word unigrams turned out to be a very competitive model with reasonable results in the tasks we have participated. An error analysis has been performed in which we identified many test examples with no linguistic evidences to distinguish among the variants.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

基于困惑度的相似语言判别方法

本文描述了Citius_Ixa_Imaxin团队提交给VarDial 2017的系统(DSL和GDI任务)。我们系统的基本策略是基于通过模型困惑计算的语言距离。我们测试过的最好的模型配置是使用单词和字符的几个n-gram模型的投票系统，即使单词unigrams在我们参与的任务中被证明是一个非常有竞争力的模型，结果也很合理。进行了错误分析，其中我们确定了许多没有语言证据的测试示例来区分变体。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Workshop on NLP for Similar Languages, Varieties and Dialects

自引率

0.00%

发文量