Multi-thread Multi-keywords Matching Approach for Uyghur Text

2013 International Conference on Asian Language Processing Pub Date : 2013-08-17 DOI:10.1109/IALP.2013.36

Xinyuan Zhao, Adili Abuliz

引用次数: 0

Abstract

Keywords matching is a preliminary means in public opinion analysis. Uyghur language is an agglutinative language, which words can be attaching by suffixes to express different semantic or syntactic in the text. Therefore, traditional matching algorithm can not be applied directly to the Uyghur text due to the Uyghur words have different surface forms in the text. In this paper, we implement a multi-keywords matching algorithm based on automaton for Uyghur text. The algorithm handles the inflection suffixes and the weakening of vowel letter in the word by use of reseverse suffixes automata and weakening of vowel restoration automata. By classification the keywords automata on the first letter of each keyword, a general multi-thread keywords matching approach for Uyghur also be proposed.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

维吾尔语文本多线程多关键词匹配方法

关键词匹配是民意分析的初步手段。维吾尔语是一种粘附性语言，词语可以通过词缀来表达文本中不同的语义或句法。因此，由于维吾尔语单词在文本中具有不同的表面形式，传统的匹配算法不能直接应用于维吾尔语文本。本文实现了一种基于自动机的维吾尔语文本多关键词匹配算法。该算法利用逆后缀自动机和弱化元音恢复自动机来处理单词中的屈折后缀和弱化元音字母。根据关键词首字母对关键词自动机进行分类，提出了一种通用的维吾尔语多线程关键词匹配方法。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2013 International Conference on Asian Language Processing

自引率

0.00%

发文量