增量分层自监督学习在设备上实现高效的无监督语音域自适应

Interspeech Pub Date : 2022-09-18 DOI:10.21437/interspeech.2022-10904

Zhouyuan Huo, DongSeon Hwang, K. Sim, Shefali Garg, Ananya Misra, Nikhil Siddhartha, Trevor Strohman, F. Beaufays

{"title":"增量分层自监督学习在设备上实现高效的无监督语音域自适应","authors":"Zhouyuan Huo, DongSeon Hwang, K. Sim, Shefali Garg, Ananya Misra, Nikhil Siddhartha, Trevor Strohman, F. Beaufays","doi":"10.21437/interspeech.2022-10904","DOIUrl":null,"url":null,"abstract":"Streaming end-to-end speech recognition models have been widely applied to mobile devices and show signiﬁcant improvement in efﬁciency. These models are typically trained on the server using transcribed speech data. However, the server data distribution can be very different from the data distribution on user devices, which could affect the model performance. There are two main challenges for on device training, limited reliable labels and limited training memory. While self-supervised learning algorithms can mitigate the mismatch between domains using unlabeled data, they are not applicable on mobile devices directly because of the memory constraint. In this paper, we propose an incremental layer-wise self-supervised learning algorithm for efﬁcient unsupervised speech domain adaptation on mobile devices, in which only one layer is updated at a time. Extensive experimental results demonstrate that the proposed algorithm achieves a 24 . 2% relative Word Error Rate (WER) improvement on the target domain compared to a supervised baseline and costs 95 . 7% less training memory than the end-to-end self-supervised learning algorithm.","PeriodicalId":73500,"journal":{"name":"Interspeech","volume":"1 1","pages":"4845-4849"},"PeriodicalIF":0.0000,"publicationDate":"2022-09-18","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":"{\"title\":\"Incremental Layer-Wise Self-Supervised Learning for Efficient Unsupervised Speech Domain Adaptation On Device\",\"authors\":\"Zhouyuan Huo, DongSeon Hwang, K. Sim, Shefali Garg, Ananya Misra, Nikhil Siddhartha, Trevor Strohman, F. Beaufays\",\"doi\":\"10.21437/interspeech.2022-10904\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Streaming end-to-end speech recognition models have been widely applied to mobile devices and show signiﬁcant improvement in efﬁciency. These models are typically trained on the server using transcribed speech data. However, the server data distribution can be very different from the data distribution on user devices, which could affect the model performance. There are two main challenges for on device training, limited reliable labels and limited training memory. While self-supervised learning algorithms can mitigate the mismatch between domains using unlabeled data, they are not applicable on mobile devices directly because of the memory constraint. In this paper, we propose an incremental layer-wise self-supervised learning algorithm for efﬁcient unsupervised speech domain adaptation on mobile devices, in which only one layer is updated at a time. Extensive experimental results demonstrate that the proposed algorithm achieves a 24 . 2% relative Word Error Rate (WER) improvement on the target domain compared to a supervised baseline and costs 95 . 7% less training memory than the end-to-end self-supervised learning algorithm.\",\"PeriodicalId\":73500,\"journal\":{\"name\":\"Interspeech\",\"volume\":\"1 1\",\"pages\":\"4845-4849\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2022-09-18\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"1\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Interspeech\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.21437/interspeech.2022-10904\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Interspeech","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.21437/interspeech.2022-10904","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 1

摘要

流式端到端语音识别模型已被广泛应用于移动设备，并显示出效率的显著提高。这些模型通常在服务器上使用转录的语音数据进行训练。然而，服务器数据分布可能与用户设备上的数据分布非常不同，这可能会影响模型性能。设备上训练有两个主要挑战，有限的可靠标签和有限的训练记忆。虽然自监督学习算法可以使用未标记的数据来减轻域之间的不匹配，但由于内存限制，它们不直接适用于移动设备。在本文中，我们提出了一种增量分层自监督学习算法，用于移动设备上有效的无监督语音域自适应，其中一次只更新一层。大量的实验结果表明，所提出的算法实现了24。与监督基线相比，目标域上2%的相对字错误率（WER）改进，并且成本为95。与端到端自监督学习算法相比，训练内存减少7%。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Incremental Layer-Wise Self-Supervised Learning for Efficient Unsupervised Speech Domain Adaptation On Device

Streaming end-to-end speech recognition models have been widely applied to mobile devices and show signiﬁcant improvement in efﬁciency. These models are typically trained on the server using transcribed speech data. However, the server data distribution can be very different from the data distribution on user devices, which could affect the model performance. There are two main challenges for on device training, limited reliable labels and limited training memory. While self-supervised learning algorithms can mitigate the mismatch between domains using unlabeled data, they are not applicable on mobile devices directly because of the memory constraint. In this paper, we propose an incremental layer-wise self-supervised learning algorithm for efﬁcient unsupervised speech domain adaptation on mobile devices, in which only one layer is updated at a time. Extensive experimental results demonstrate that the proposed algorithm achieves a 24 . 2% relative Word Error Rate (WER) improvement on the target domain compared to a supervised baseline and costs 95 . 7% less training memory than the end-to-end self-supervised learning algorithm.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Interspeech

自引率

0.00%

发文量