Audiovisual to area and length functions inversion of human vocal tract

2014 22nd European Signal Processing Conference (EUSIPCO) Pub Date : 2014-11-13 DOI:10.5281/ZENODO.43890

Benjamin Elie, Y. Laprie

引用次数: 5

Abstract

This paper proposes a multimodal approach to estimate the area function and the length of the vocal tract of oral vowels. The method is based on an iterative technique consisting in deforming an initial area function so that the output acoustic vector matches a specified target. The chosen acoustic vector is the formant frequency pattern. In order to regularize the ill-problem, several constraints are added to the algorithm. First, the lip termination area is estimated via a facial capture software. Then, the area function is constrained in such a way that it does not get too far from a neutral position, and it does not change too quickly from a temporal frame to the next, when dealing with dynamic inversion. The method proves to be efficient to approximate the area function and the length of the vocal tract for oral french vowels, both in static and dynamic configurations.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

声道面积和长度函数倒置的视听效果

本文提出了一种多模态方法来估计口腔元音的面积函数和声道长度。该方法基于一种迭代技术，包括变形初始面积函数，使输出声矢量与指定目标匹配。所选择的声矢量是形成峰频率模式。为了使病态问题正则化，在算法中加入了若干约束条件。首先，通过面部捕捉软件估计唇终止区域。然后，在处理动态反演时，对面积函数进行约束，使其不会离中性位置太远，也不会从一个时间帧到下一个时间帧变化太快。该方法在静态和动态情况下都能有效地逼近法语元音的面积函数和声道长度。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2014 22nd European Signal Processing Conference (EUSIPCO)

自引率

0.00%

发文量