Integration of articulatory knowledge and voicing features based on DNN/HMM for Mandarin speech recognition

2015 International Joint Conference on Neural Networks (IJCNN) Pub Date : 2015-07-12 DOI:10.1109/IJCNN.2015.7280396

Ying-Wei Tan, Wenju Liu, Wei Jiang, Hao Zheng

{"title":"Integration of articulatory knowledge and voicing features based on DNN/HMM for Mandarin speech recognition","authors":"Ying-Wei Tan, Wenju Liu, Wei Jiang, Hao Zheng","doi":"10.1109/IJCNN.2015.7280396","DOIUrl":null,"url":null,"abstract":"Speech production knowledge has been used to enhance the phonetic representation and the performance of automatic speech recognition (ASR) systems successfully. Representations of speech production make simple explanations for many phenomena observed in speech. These phenomena can not be easily analyzed from either acoustic signal or phonetic transcription alone. One of the most important aspects of speech production knowledge is the use of articulatory knowledge, which describes the smooth and continuous movements in the vocal tract. In this paper, we present a new articulatory model to provide available information for rescoring the speech recognition lattice hypothesis. The articulatory model consists of a feature front-end, which computes a voicing feature based on a spectral harmonics correlation (SHC) function, and a back-end based on the combination of deep neural networks (DNNs) and hidden Markov models (HMMs). The voicing features are incorporated with standard Mel frequency cepstral coefficients (MFCCs) using heteroscedastic linear discriminant analysis (HLDA) to compensate the speech recognition accuracy rates. Moreover, the advantages of two different models are taken into account by the algorithm, which retains deep learning properties of DNNs, while modeling the articulatory context powerfully through HMMs. Mandarin speech recognition experiments show the proposed method achieves significant improvements in speech recognition performance over the system using MFCCs alone.","PeriodicalId":6539,"journal":{"name":"2015 International Joint Conference on Neural Networks (IJCNN)","volume":"50 1","pages":"1-8"},"PeriodicalIF":0.0000,"publicationDate":"2015-07-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"3","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2015 International Joint Conference on Neural Networks (IJCNN)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/IJCNN.2015.7280396","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 3

Abstract

Speech production knowledge has been used to enhance the phonetic representation and the performance of automatic speech recognition (ASR) systems successfully. Representations of speech production make simple explanations for many phenomena observed in speech. These phenomena can not be easily analyzed from either acoustic signal or phonetic transcription alone. One of the most important aspects of speech production knowledge is the use of articulatory knowledge, which describes the smooth and continuous movements in the vocal tract. In this paper, we present a new articulatory model to provide available information for rescoring the speech recognition lattice hypothesis. The articulatory model consists of a feature front-end, which computes a voicing feature based on a spectral harmonics correlation (SHC) function, and a back-end based on the combination of deep neural networks (DNNs) and hidden Markov models (HMMs). The voicing features are incorporated with standard Mel frequency cepstral coefficients (MFCCs) using heteroscedastic linear discriminant analysis (HLDA) to compensate the speech recognition accuracy rates. Moreover, the advantages of two different models are taken into account by the algorithm, which retains deep learning properties of DNNs, while modeling the articulatory context powerfully through HMMs. Mandarin speech recognition experiments show the proposed method achieves significant improvements in speech recognition performance over the system using MFCCs alone.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

基于深度神经网络/HMM的普通话语音识别中发音知识与语音特征的整合

语音生成知识已被成功地应用于语音自动识别系统的语音表示和性能提升。言语产生表征对言语中观察到的许多现象作出了简单的解释。这些现象单从声学信号或音标分析都不容易。语音产生知识的一个最重要的方面是发音知识的使用，它描述了声道中流畅和连续的运动。在本文中，我们提出了一个新的发音模型，为重新记录语音识别晶格假设提供了可用的信息。该发音模型包括基于谱谐波相关(SHC)函数计算语音特征的特征前端和基于深度神经网络(dnn)和隐马尔可夫模型(hmm)相结合的后端。利用异方差线性判别分析(HLDA)将语音特征与标准Mel频率倒谱系数(MFCCs)相结合，对语音识别准确率进行补偿。此外，该算法考虑了两种不同模型的优点，保留了dnn的深度学习特性，同时通过hmm对发音上下文进行了强大的建模。普通话语音识别实验表明，与单独使用mfc的系统相比，该方法在语音识别性能上有显著提高。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2015 International Joint Conference on Neural Networks (IJCNN)

自引率

0.00%

发文量