Model-based dereverberation in the logmelspec domain for robust distant-talking speech recognition

2010 IEEE International Conference on Acoustics, Speech and Signal Processing Pub Date : 2010-06-28 DOI:10.1109/ICASSP.2010.5495671

A. Sehr, R. Maas, Walter Kellermann

引用次数: 18

Abstract

The REMOS (REverberation MOdeling for Speech recognition) concept for reverberation-robust distant-talking speech recognition, introduced in [1] for melspectral features, is extended in this contribution to logarithmic melspectral (logmelspec) features. Based on a combined acoustic model consisting of a hidden Markov model network and a reverberation model, REMOS determines clean-speech and reverberation estimates during recognition by an inner optimization operation. A reformulation of this inner optimization problem for logmelspec features, allowing an efficient solution by nonlinear optimization algorithms, is derived in this paper so that an efficient implementation of REMOS for logmelspec features becomes possible. Connected digit recognition experiments show that the proposed REMOS implementation significantly outperforms reverberantly-trained HMMs in highly reverberant environments.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

基于logmelspec域模型的鲁棒远距离语音识别去噪

[1]中引入的用于混响鲁棒远距离语音识别的REMOS (REverberation MOdeling for Speech recognition)概念用于melspectrum特征，在本贡献中扩展到对数melspectrum (logmelspec)特征。REMOS基于由隐马尔可夫模型网络和混响模型组成的组合声学模型，通过内部优化操作确定识别过程中的干净语音和混响估计。本文推导了logmelspec特征的内部优化问题的一个重新表述，允许通过非线性优化算法进行有效的解决，从而使logmelspec特征的REMOS的有效实现成为可能。连接数字识别实验表明，所提出的REMOS实现在高混响环境下显著优于混响训练的hmm。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2010 IEEE International Conference on Acoustics, Speech and Signal Processing

自引率

0.00%

发文量

期刊最新文献

Interactive tone mapping for High Dynamic Range video Search error risk minimization in Viterbi beam search for speech recognition Predicting interruptions in dyadic spoken interactions Simple methods for improving speaker-similarity of HMM-based speech synthesis Model-based dereverberation in the logmelspec domain for robust distant-talking speech recognition