{"title":"Bounded support Gaussian mixture modeling of speech spectra","authors":"J. Lindblom, J. Samuelsson","doi":"10.1109/TSA.2002.805639","DOIUrl":null,"url":null,"abstract":"Lately, Gaussian mixture (GM) models have found new applications in speech processing, and particularly in speech coding. This paper provides a review of GM based quantization and prediction. The main contribution is a discussion on GM model optimization. Two previously presented algorithms of EM-type are analyzed in some detail, and models are estimated and evaluated experimentally using theoretical measures as well as GM based speech spectrum coding and prediction. It has been argued that since many sources have a bounded support, this should be utilized in both the choice of model, and the optimization algorithm. By low-dimensional modeling examples, illustrating the behavior of the two algorithms graphically, and by full-scale evaluation of GM based systems, the advantages of a bounded support approach are quantified. For all evaluation techniques in the study, model accuracy is improved when the bounded support approach is adopted. The gains are typically largest for models with diagonal covariance matrices.","PeriodicalId":13155,"journal":{"name":"IEEE Trans. Speech Audio Process.","volume":"9 1","pages":"88-99"},"PeriodicalIF":0.0000,"publicationDate":"2003-02-19","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"66","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"IEEE Trans. Speech Audio Process.","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/TSA.2002.805639","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 66
Abstract
Lately, Gaussian mixture (GM) models have found new applications in speech processing, and particularly in speech coding. This paper provides a review of GM based quantization and prediction. The main contribution is a discussion on GM model optimization. Two previously presented algorithms of EM-type are analyzed in some detail, and models are estimated and evaluated experimentally using theoretical measures as well as GM based speech spectrum coding and prediction. It has been argued that since many sources have a bounded support, this should be utilized in both the choice of model, and the optimization algorithm. By low-dimensional modeling examples, illustrating the behavior of the two algorithms graphically, and by full-scale evaluation of GM based systems, the advantages of a bounded support approach are quantified. For all evaluation techniques in the study, model accuracy is improved when the bounded support approach is adopted. The gains are typically largest for models with diagonal covariance matrices.