MICHAEL: Mining Character-level Patterns for Arabic Dialect Identification (MADAR Challenge)

WANLP@ACL 2019 Pub Date : 2019-08-01 DOI:10.18653/v1/W19-4627

Dhaou Ghoul, Gaël Lejeune

引用次数: 5

Abstract

We present MICHAEL, a simple lightweight method for automatic Arabic Dialect Identification on the MADAR travel domain Dialect Identification (DID). MICHAEL uses simple character-level features in order to perform a pre-processing free classification. More precisely, Character N-grams extracted from the original sentences are used to train a Multinomial Naive Bayes classifier. This system achieved an official score (accuracy) of 53.25% with 1<=N<=3 but showed a much better result with character 4-grams (62.17% accuracy).

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

MICHAEL:挖掘阿拉伯语方言识别的字符级模式(MADAR挑战)

本文提出了一种基于MADAR旅行域方言识别(DID)的简易轻量级阿拉伯语方言自动识别方法MICHAEL。MICHAEL使用简单的字符级特征来执行预处理自由分类。更准确地说，从原始句子中提取的字符N-grams用于训练多项式朴素贝叶斯分类器。该系统在1<=N<=3时的官方得分(正确率)为53.25%，但在4克字符时的结果要好得多(正确率为62.17%)。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

WANLP@ACL 2019

自引率

0.00%

发文量