基于分解成分词声学模型的情绪语音识别

2013 2nd IAPR Asian Conference on Pattern Recognition Pub Date : 2013-11-05 DOI:10.1109/ACPR.2013.13

Vivatchai Kaveeta, K. Patanukhom

{"title":"基于分解成分词声学模型的情绪语音识别","authors":"Vivatchai Kaveeta, K. Patanukhom","doi":"10.1109/ACPR.2013.13","DOIUrl":null,"url":null,"abstract":"This paper presents a novel approach for emotional speech recognition. Instead of using a full length of speech for classification, the proposed method decomposes speech signals into component words, groups the words into segments and generates an acoustic model for each segment by using features such as audio power, MFCC, log attack time, spectrum spread and segment duration. Based on the proposed segment-based classification, unknown speech signals can be recognized into sequences of segment emotions. Emotion profiles (EPs) are extracted from the emotion sequences. Finally, speech emotion can be determined by using EP as features. Experiments are conducted by using 6,810 training samples and 722 test samples which are composed of eight emotional classes from IEMOCAP database. In comparison with a conventional method, the proposed method can improve recognition rate from 46.81% to 58.59% in eight emotion classification and from 60.18% to 71.25% in four emotion classification.","PeriodicalId":365633,"journal":{"name":"2013 2nd IAPR Asian Conference on Pattern Recognition","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2013-11-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Emotional Speech Recognition Using Acoustic Models of Decomposed Component Words\",\"authors\":\"Vivatchai Kaveeta, K. Patanukhom\",\"doi\":\"10.1109/ACPR.2013.13\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"This paper presents a novel approach for emotional speech recognition. Instead of using a full length of speech for classification, the proposed method decomposes speech signals into component words, groups the words into segments and generates an acoustic model for each segment by using features such as audio power, MFCC, log attack time, spectrum spread and segment duration. Based on the proposed segment-based classification, unknown speech signals can be recognized into sequences of segment emotions. Emotion profiles (EPs) are extracted from the emotion sequences. Finally, speech emotion can be determined by using EP as features. Experiments are conducted by using 6,810 training samples and 722 test samples which are composed of eight emotional classes from IEMOCAP database. In comparison with a conventional method, the proposed method can improve recognition rate from 46.81% to 58.59% in eight emotion classification and from 60.18% to 71.25% in four emotion classification.\",\"PeriodicalId\":365633,\"journal\":{\"name\":\"2013 2nd IAPR Asian Conference on Pattern Recognition\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2013-11-05\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2013 2nd IAPR Asian Conference on Pattern Recognition\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ACPR.2013.13\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2013 2nd IAPR Asian Conference on Pattern Recognition","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ACPR.2013.13","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

提出了一种新的情感语音识别方法。该方法不使用完整的语音长度进行分类，而是将语音信号分解为成分词，将词分组为段，并利用音频功率、MFCC、日志攻击时间、频谱扩展和段持续时间等特征为每个段生成声学模型。基于所提出的基于片段的分类方法，可以将未知语音信号识别为片段情绪序列。从情感序列中提取情感轮廓(EPs)。最后，利用EP作为特征来确定语音情绪。实验采用IEMOCAP数据库中的6810个训练样本和722个测试样本组成的8个情绪类。与传统方法相比，该方法在8种情绪分类中将识别率从46.81%提高到58.59%，在4种情绪分类中将识别率从60.18%提高到71.25%。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Emotional Speech Recognition Using Acoustic Models of Decomposed Component Words

This paper presents a novel approach for emotional speech recognition. Instead of using a full length of speech for classification, the proposed method decomposes speech signals into component words, groups the words into segments and generates an acoustic model for each segment by using features such as audio power, MFCC, log attack time, spectrum spread and segment duration. Based on the proposed segment-based classification, unknown speech signals can be recognized into sequences of segment emotions. Emotion profiles (EPs) are extracted from the emotion sequences. Finally, speech emotion can be determined by using EP as features. Experiments are conducted by using 6,810 training samples and 722 test samples which are composed of eight emotional classes from IEMOCAP database. In comparison with a conventional method, the proposed method can improve recognition rate from 46.81% to 58.59% in eight emotion classification and from 60.18% to 71.25% in four emotion classification.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2013 2nd IAPR Asian Conference on Pattern Recognition

自引率

0.00%

发文量

期刊最新文献

Automatic Compensation of Radial Distortion by Minimizing Entropy of Histogram of Oriented Gradients A Robust and Efficient Minutia-Based Fingerprint Matching Algorithm Sclera Recognition - A Survey A Non-local Sparse Model for Intrinsic Images Classification Based on Boolean Algebra and Its Application to the Prediction of Recurrence of Liver Cancer