Factor analysis based session variability compensation for Automatic Speech Recognition

2011 IEEE Workshop on Automatic Speech Recognition & Understanding Pub Date : 2011-12-01 DOI:10.1109/ASRU.2011.6163920

Mickael Rouvier, M. Bouallegue, D. Matrouf, G. Linarès

引用次数: 3

Abstract

In this paper we propose a new feature normalization based on Factor Analysis (FA) for the problem of acoustic variability in Automatic Speech Recognition (ASR). The FA paradigm was previously used in the field of ASR, in order to model the usefull information: the HMM state dependent acoustic information. In this paper, we propose to use the FA paradigm to model the useless information (speaker- or channel-variability) in order to remove it from acoustic data frames. The transformed training data frames are then used to train new HMM models using the standard training algorithm. The transformation is also applied to the test data before the decoding process. With this approach we obtain, on french broadcast news, an absolute WER reduction of 1.3%.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

基于因子分析的会话可变性自动语音识别补偿

针对自动语音识别中的声学变异性问题，提出了一种基于因子分析的特征归一化方法。先前在ASR领域中使用了FA范式，以建模有用的信息:HMM状态相关的声学信息。在本文中，我们建议使用FA范式对无用信息(说话者或信道可变性)进行建模，以便从声学数据帧中删除无用信息。然后使用转换后的训练数据帧使用标准训练算法训练新的HMM模型。在解码之前，还对测试数据进行了转换。通过这种方法，我们在法国广播新闻中获得了绝对减少1.3%的WER。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2011 IEEE Workshop on Automatic Speech Recognition & Understanding

自引率

0.00%

发文量

期刊最新文献

Applying feature bagging for more accurate and robust automated speaking assessment Towards choosing better primes for spoken dialog systems Accent level adjustment in bilingual Thai-English text-to-speech synthesis Fast speaker diarization using a high-level scripting language Evaluating prosodic features for automated scoring of non-native read speech