A unified trajectory tiling approach to high quality TTS and cross-lingual voice transformation

2012 8th International Symposium on Chinese Spoken Language Processing Pub Date : 2012-12-01 DOI:10.1109/ISCSLP.2012.6423506

Yao Qian, F. Soong

引用次数: 0

Abstract

In human-machine speech communication, it is technically challenging to make the machine talk as naturally as human so as to facilitate “frictionless” interactions, or make a human user to feel the communication is as natural as human-human. We propose a trajectory tiling approach to high quality speech synthesis, where the speech parameter trajectories, extracted from natural, processed, or synthesized speech, are used to guide the search for the best sequence of waveform segment “tiles” stored in a pre-recorded speech database. We test our approach in both TTS and cross-lingual voice transformation applications. Experimental results show that the proposed trajectory tiling approach can render speech which is both natural and highly intelligible. The perceived high quality speech is also confirmed in objective and subjective tests.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

统一轨迹平铺方法实现高质量TTS和跨语言语音转换

在人机语音交流中，如何让机器像人一样自然地说话，从而促进“无摩擦”的交互，或者让人类用户感觉交流像人与人一样自然，在技术上是一个挑战。我们提出了一种用于高质量语音合成的轨迹平铺方法，其中从自然、处理或合成语音中提取的语音参数轨迹用于指导搜索存储在预录制语音数据库中的最佳波形段“平铺”序列。我们在TTS和跨语言语音转换应用程序中测试了我们的方法。实验结果表明，所提出的轨迹平铺方法能够呈现出既自然又高可理解的语音。感知到的高质量语音也在客观和主观测试中得到证实。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2012 8th International Symposium on Chinese Spoken Language Processing

自引率

0.00%

发文量