Analysis of acoustic-phonetic variations in fluent speech using TIMIT

1995 International Conference on Acoustics, Speech, and Signal Processing Pub Date : 1995-05-09 DOI:10.1109/ICASSP.1995.479399

Don X. Sun, L. Deng

引用次数: 18

Abstract

We propose a hierarchically structured analysis of variance (ANOVA) method to analyze, in a quantitative manner, the contributions of various identifiable factors to the overall acoustic variability exhibited in fluent speech data of TIMIT processed in the form of mel-frequency cepstral coefficients. The results of the analysis show that the greatest acoustic variability in TIMIT data is explained by the difference among distinct phonetic labels in TIMIT, followed by the phonetic context difference given a fixed phonetic label. The variability among sequential sub-segments within each TIMIT-defined phonetic segment is found to be significantly greater than the gender, dialect region, and speaker factors. Our results serve to provide useful insights to the understanding of the roles of various components of speech recognizers in contributing to the ultimate speech recognition performance.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

用TIMIT分析流利言语的声音变化

我们提出了一种层次结构的方差分析(ANOVA)方法，以定量的方式分析各种可识别因素对以梅尔频率倒谱系数形式处理的TIMIT流畅语音数据中所表现出的整体声学变异性的贡献。分析结果表明，TIMIT数据中最大的声学变异性是由TIMIT中不同语音标签之间的差异造成的，其次是固定语音标签下的语音语境差异。在每个由timit定义的语音段中，顺序子段之间的变异性明显大于性别、方言区域和说话人因素。我们的研究结果为理解语音识别器的各个组成部分在最终语音识别性能中的作用提供了有用的见解。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

1995 International Conference on Acoustics, Speech, and Signal Processing

自引率

0.00%

发文量

期刊最新文献

Language identification with phonological and lexical models Computationally efficient wavelet packet coding of wide-band stereo audio signals Signaling techniques using solitons Blind source detection and separation using second order non-stationarity On blind channel identification for impulsive signal environments