语音信号自动识别系统对语音信号退化因素的鲁棒性分析

2015 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA) Pub Date : 2015-12-28 DOI:10.1109/SPA.2015.7365136

J. Oska, J. Wojtun, K. Wodecki, Z. Piotrowski

{"title":"语音信号自动识别系统对语音信号退化因素的鲁棒性分析","authors":"J. Oska, J. Wojtun, K. Wodecki, Z. Piotrowski","doi":"10.1109/SPA.2015.7365136","DOIUrl":null,"url":null,"abstract":"In the article there are presented the results of research on the influence of the lossy compression, used in codecs G.711, G.723.1 and iLBC, on the efficiency of isolated speech phrase recognition. In the research the degree of robustness against degrading factors in the parameterisation method of audio signal LPCC and MFCC (Linear Prediction Cepstral Coefficients, Mel Frequency Cepstral Coefficients) is compared. The research is based on the classifier of improved Gaussian mixtures making allowance for Universal Background Model GMM-UBM (Gaussian Mixtures Model - Universal Background Model). The research was conducted on the database composed of 3000 isolated speech phrases.","PeriodicalId":423880,"journal":{"name":"2015 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA)","volume":"5 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2015-12-28","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Robustness analysis of automatic speech signal recognition system against factors degrading speech signal\",\"authors\":\"J. Oska, J. Wojtun, K. Wodecki, Z. Piotrowski\",\"doi\":\"10.1109/SPA.2015.7365136\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In the article there are presented the results of research on the influence of the lossy compression, used in codecs G.711, G.723.1 and iLBC, on the efficiency of isolated speech phrase recognition. In the research the degree of robustness against degrading factors in the parameterisation method of audio signal LPCC and MFCC (Linear Prediction Cepstral Coefficients, Mel Frequency Cepstral Coefficients) is compared. The research is based on the classifier of improved Gaussian mixtures making allowance for Universal Background Model GMM-UBM (Gaussian Mixtures Model - Universal Background Model). The research was conducted on the database composed of 3000 isolated speech phrases.\",\"PeriodicalId\":423880,\"journal\":{\"name\":\"2015 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA)\",\"volume\":\"5 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2015-12-28\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2015 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/SPA.2015.7365136\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2015 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/SPA.2015.7365136","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

本文给出了在G.711、G.723.1和iLBC编解码器中使用有损压缩对孤立语音短语识别效率影响的研究结果。在研究中比较了音频信号参数化方法LPCC和MFCC(线性预测倒谱系数，Mel频率倒谱系数)对退化因子的鲁棒性。本研究基于改进的高斯混合分类器，并考虑通用背景模型GMM-UBM(高斯混合模型-通用背景模型)。该研究是在由3000个孤立的语音短语组成的数据库上进行的。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Robustness analysis of automatic speech signal recognition system against factors degrading speech signal

In the article there are presented the results of research on the influence of the lossy compression, used in codecs G.711, G.723.1 and iLBC, on the efficiency of isolated speech phrase recognition. In the research the degree of robustness against degrading factors in the parameterisation method of audio signal LPCC and MFCC (Linear Prediction Cepstral Coefficients, Mel Frequency Cepstral Coefficients) is compared. The research is based on the classifier of improved Gaussian mixtures making allowance for Universal Background Model GMM-UBM (Gaussian Mixtures Model - Universal Background Model). The research was conducted on the database composed of 3000 isolated speech phrases.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2015 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA)

自引率

0.00%

发文量

期刊最新文献

Influence of simultaneous spoken sentences on the properties of spectral peaks Measurements and visualization of sound field distribution around organ pipe Representing the evolving temporal envelope of musical instruments sounds using Computer Vision methods Irregular sampling for X-ray imaging simulation An enhancement of software metrics as failure predictors