基于感知mvdr的哈萨克语语音识别的无监督内置说话人归一化

2014 IEEE 8th International Conference on Application of Information and Communication Technologies (AICT) Pub Date : 2014-10-01 DOI:10.1109/ICAICT.2014.7035914

Zhandos Yessenbayev, U. Yapanel

{"title":"基于感知mvdr的哈萨克语语音识别的无监督内置说话人归一化","authors":"Zhandos Yessenbayev, U. Yapanel","doi":"10.1109/ICAICT.2014.7035914","DOIUrl":null,"url":null,"abstract":"In this work we present a novel approach to unsupervised speaker normalization on top of the Perceptual MVDR-based Built-in Speaker Normalization technique. We showed that the proposed method can be efficient for the task of phonetic recognition on TIMIT and then applied it to Kazakh speech recognition. From the experiments, we see that this method is able to improve the relative performance of ASR systems up to 20% The analysis of the optimal warp factor selection by the algorithm revealed a nice gender separation ability which may be used for gender/speaker classification tasks.","PeriodicalId":103329,"journal":{"name":"2014 IEEE 8th International Conference on Application of Information and Communication Technologies (AICT)","volume":"24 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2014-10-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Perceptual MVDR-based unsupervised built-in speaker normalization for Kazakh speech recognition\",\"authors\":\"Zhandos Yessenbayev, U. Yapanel\",\"doi\":\"10.1109/ICAICT.2014.7035914\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In this work we present a novel approach to unsupervised speaker normalization on top of the Perceptual MVDR-based Built-in Speaker Normalization technique. We showed that the proposed method can be efficient for the task of phonetic recognition on TIMIT and then applied it to Kazakh speech recognition. From the experiments, we see that this method is able to improve the relative performance of ASR systems up to 20% The analysis of the optimal warp factor selection by the algorithm revealed a nice gender separation ability which may be used for gender/speaker classification tasks.\",\"PeriodicalId\":103329,\"journal\":{\"name\":\"2014 IEEE 8th International Conference on Application of Information and Communication Technologies (AICT)\",\"volume\":\"24 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2014-10-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2014 IEEE 8th International Conference on Application of Information and Communication Technologies (AICT)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ICAICT.2014.7035914\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2014 IEEE 8th International Conference on Application of Information and Communication Technologies (AICT)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICAICT.2014.7035914","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

在这项工作中，我们在基于感知mvdr的内置说话人归一化技术的基础上提出了一种新的无监督说话人归一化方法。结果表明，该方法可以有效地完成语音识别任务，并将其应用于哈萨克语语音识别中。实验结果表明，该方法可将ASR系统的相对性能提高20%以上。通过对算法的最优扭曲因子选择的分析，表明该算法具有良好的性别分离能力，可用于性别/说话人分类任务。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Perceptual MVDR-based unsupervised built-in speaker normalization for Kazakh speech recognition

In this work we present a novel approach to unsupervised speaker normalization on top of the Perceptual MVDR-based Built-in Speaker Normalization technique. We showed that the proposed method can be efficient for the task of phonetic recognition on TIMIT and then applied it to Kazakh speech recognition. From the experiments, we see that this method is able to improve the relative performance of ASR systems up to 20% The analysis of the optimal warp factor selection by the algorithm revealed a nice gender separation ability which may be used for gender/speaker classification tasks.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2014 IEEE 8th International Conference on Application of Information and Communication Technologies (AICT)

自引率

0.00%

发文量

期刊最新文献

A new robust binary image embedding algorithm in discrete wavelet domain Polyalphabetic Euclidean ciphers Complex system state generalized presentation based on concepts Using a knowledge base in developing modification for MS Dynamics AX TOFI technology capabilities for data processing and visualization