A mixed-excitation frequency domain model for time-scale pitch-scale modification of speech

5th International Conference on Spoken Language Processing (ICSLP 1998) Pub Date : 1998-11-30 DOI:10.21437/ICSLP.1998-16

A. Acero

引用次数: 3

Abstract

This paper presents a time-scale pitch-scale modification technique for concatenative speech synthesis. The method is based on a frequency domain source-filter model, where the source is modeled as a mixed excitation. This model is highly coupled with a compression scheme that result in compact acoustic inventories. When compared to the approach in the Whistler system using no mixed excitation, the new method shows improvement in voiced fricatives and over-stretched voiced sounds. In addition, it allows for spectral manipulation such as smoothing of discontinuities at unit boundaries, voice transformations or loudness equalization.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

语音时阶音高阶修正的混合激励频域模型

本文提出了一种用于串联语音合成的时尺度音阶修正技术。该方法基于频域源-滤波器模型，其中源被建模为混合激励。该模型与压缩方案高度耦合，从而产生紧凑的声学清单。与没有混合激励的Whistler系统的方法相比，新方法在浊音摩擦音和过伸浊音方面表现出改善。此外，它允许频谱操作，如平滑不连续在单位边界，语音转换或响度均衡。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

5th International Conference on Spoken Language Processing (ICSLP 1998)

自引率

0.00%

发文量

期刊最新文献

Assimilation of place in Japanese and dutch Articulatory analysis using a codebook for articulatory based low bit-rate speech coding Phonetic and phonological characteristics of paralinguistic information in spoken Japanese HMM-based visual speech recognition using intensity and location normalization Speech recognition via phonetically featured syllables