Blind Extraction of Moving Audio Source in a Challenging Environment Supported by Speaker Identification Via X-Vectors

ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) Pub Date : 2021-06-06 DOI:10.1109/ICASSP39728.2021.9414331

J. Málek, Jakub Janský, Tomás Kounovský, Zbyněk Koldovský, J. Zdánský

{"title":"Blind Extraction of Moving Audio Source in a Challenging Environment Supported by Speaker Identification Via X-Vectors","authors":"J. Málek, Jakub Janský, Tomás Kounovský, Zbyněk Koldovský, J. Zdánský","doi":"10.1109/ICASSP39728.2021.9414331","DOIUrl":null,"url":null,"abstract":"We propose a novel approach for semi-supervised extraction of a moving audio source of interest (SOI) applicable in reverberant and noisy environments. The blind part of the method is based on independent vector extraction (IVE) and uses the recently proposed constant separating vector (CSV) mixing model. This model allows for changes of mixing parameters within the processed interval of the mixture, which potentially leads to higher accuracy of SOI estimation. The supervised part of the method concerns a pilot signal, which is related to the SOI and ensures the convergence of the blind method towards the SOI. The pilot is based on robust detection of frames where SOI is dominant via speaker embeddings called X-vectors. Robustness of the detection is achieved through augmentation of the data for the supervised training of the X-vectors. The pilot-supported extraction yields significantly better performance compared to its unsupervised counterpart identifying SOI solely using the initialization.","PeriodicalId":347060,"journal":{"name":"ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2021-06-06","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"6","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICASSP39728.2021.9414331","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 6

Abstract

We propose a novel approach for semi-supervised extraction of a moving audio source of interest (SOI) applicable in reverberant and noisy environments. The blind part of the method is based on independent vector extraction (IVE) and uses the recently proposed constant separating vector (CSV) mixing model. This model allows for changes of mixing parameters within the processed interval of the mixture, which potentially leads to higher accuracy of SOI estimation. The supervised part of the method concerns a pilot signal, which is related to the SOI and ensures the convergence of the blind method towards the SOI. The pilot is based on robust detection of frames where SOI is dominant via speaker embeddings called X-vectors. Robustness of the detection is achieved through augmentation of the data for the supervised training of the X-vectors. The pilot-supported extraction yields significantly better performance compared to its unsupervised counterpart identifying SOI solely using the initialization.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

基于x向量说话人识别的环境下运动声源的盲提取

我们提出了一种适用于混响和噪声环境的移动音频感兴趣源(SOI)的半监督提取方法。该方法的盲区部分基于独立向量提取(IVE)，并使用了最近提出的恒定分离向量(CSV)混合模型。该模型允许在混合物的处理区间内混合参数的变化，这可能导致更高的SOI估计精度。该方法的监督部分涉及与SOI相关的导频信号，保证了盲法对SOI的收敛性。该导频是基于通过称为x向量的扬声器嵌入对SOI占主导地位的帧进行鲁棒检测。检测的鲁棒性是通过增强x向量的监督训练数据来实现的。与仅使用初始化来识别SOI的无监督提取相比，pilot支持的提取产生了明显更好的性能。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

自引率

0.00%

发文量