Latent dirichlet language model for speech recognition

2008 IEEE Spoken Language Technology Workshop Pub Date : 2008-12-01 DOI:10.1109/SLT.2008.4777875

Jen-Tzung Chien, C. Chueh

引用次数: 32

Abstract

Latent Dirichlet allocation (LDA) has been successfully presented for document modeling and classification. LDA calculates the document probability based on bag-of-words scheme without considering the sequence of words. This model discovers the topic structure at document level, which is different from the concern of word prediction in speech recognition. In this paper, we present a new latent Dirichlet language model (LDLM) for modeling of word sequence. A new Bayesian framework is introduced by merging the Dirichlet priors to characterize the uncertainty of latent topics of n-gram events. The robust topic-based language model is established accordingly. In the experiments, we implement LDLM for continuous speech recognition and obtain better performance than probabilistic latent semantic analysis (PLSA) based language method.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

语音识别的潜在狄利克雷语言模型

潜在狄利克雷分配(Latent Dirichlet allocation, LDA)已被成功地用于文档建模和分类。LDA在不考虑词序列的情况下，基于词袋方案计算文档概率。该模型在文档层面发现主题结构，不同于语音识别中对词预测的关注。本文提出了一种新的用于词序列建模的潜在狄利克雷语言模型(LDLM)。引入了一种新的贝叶斯框架，通过合并Dirichlet先验来表征n-gram事件潜在主题的不确定性。建立了基于主题的鲁棒语言模型。在实验中，我们实现了LDLM用于连续语音识别，并获得了比基于概率潜在语义分析(PLSA)的语言方法更好的性能。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2008 IEEE Spoken Language Technology Workshop

自引率

0.00%

发文量

期刊最新文献

“Who is this” quiz dialogue system and users' evaluation Latent dirichlet language model for speech recognition Modelling user behaviour in the HIS-POMDP dialogue manager A syntactic language model based on incremental CCG parsing Improving word segmentation for Thai speech translation