PDTSL: An annotated resource for speech reconstruction

2008 IEEE Spoken Language Technology Workshop Pub Date : 2008-12-01 DOI:10.1109/SLT.2008.4777848

Jan Hajic, Silvie Cinková, Marie Mikulová, P. Pajas, J. Ptáček, J. Toman, Zdenka Uresová

引用次数: 13

Abstract

We present a description of a new resource (Prague Dependency Treebank of Spoken Language) being created for English and Czech to be used for the task of speech understanding, broad natural language analysis for dialog systems and other speech-related tasks, including speech editing. The resources we have created so far contain audio and a standard transcription of spontaneous speech, but as a novel layer, we add an edited (ldquoreconstructedrdquo) version of the spoken utterances. These edits go beyond the scope of current speech reconstruction efforts in that we allow, on top of the usual deletions of speech artifacts, fillers, etc. also for word modifications, insertions and word order changes. We have used both monologue and dialogue recordings in English and Czech to verify the feasibility of such transcription. We have also assessed the quality of the resulting annotation since the relative freedom of the editing raises an issue of what a ldquocorrectrdquo annotation is.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

PDTSL:语音重建的带注释资源

我们介绍了一个为英语和捷克语创建的新资源(Prague Dependency Treebank of Spoken Language)的描述，用于语音理解任务、对话系统的广泛自然语言分析和其他语音相关任务，包括语音编辑。到目前为止，我们创建的资源包含音频和自发语音的标准转录，但作为一个新颖的层，我们添加了语音的编辑(ldquoreconstructedquo)版本。这些编辑超出了当前语音重建工作的范围，因为我们允许，除了通常的语音工件删除，填充等之外，还允许修改，插入和词序更改。我们使用了英语和捷克语的独白和对话录音来验证这种转录的可行性。我们还评估了最终注释的质量，因为编辑的相对自由提出了一个问题，即什么是最不正确的注释。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2008 IEEE Spoken Language Technology Workshop

自引率

0.00%

发文量

期刊最新文献

“Who is this” quiz dialogue system and users' evaluation Latent dirichlet language model for speech recognition Modelling user behaviour in the HIS-POMDP dialogue manager A syntactic language model based on incremental CCG parsing Improving word segmentation for Thai speech translation