从语音到音乐的迁移学习:对语言敏感的情感识别模型

2020 28th European Signal Processing Conference (EUSIPCO) Pub Date : 2021-01-24 DOI:10.23919/Eusipco47968.2020.9287548

Juan Sebastián Gómez Cañón, Estefanía Cano, P. Herrera, E. Gómez

{"title":"从语音到音乐的迁移学习:对语言敏感的情感识别模型","authors":"Juan Sebastián Gómez Cañón, Estefanía Cano, P. Herrera, E. Gómez","doi":"10.23919/Eusipco47968.2020.9287548","DOIUrl":null,"url":null,"abstract":"In this study, we address emotion recognition using unsupervised feature learning from speech data, and test its transferability to music. Our approach is to pre-train models using speech in English and Mandarin, and then fine-tune them with excerpts of music labeled with categories of emotion. Our initial hypothesis is that features automatically learned from speech should be transferable to music. Namely, we expect the intra-linguistic setting (e.g., pre-training on speech in English and fine-tuning on music in English) should result in improved performance over the cross-linguistic setting (e.g., pre-training on speech in English and fine-tuning on music in Mandarin). Our results confirm previous research on cross-domain transferability, and encourage research towards language-sensitive Music Emotion Recognition (MER) models.","PeriodicalId":6705,"journal":{"name":"2020 28th European Signal Processing Conference (EUSIPCO)","volume":"50 1","pages":"136-140"},"PeriodicalIF":0.0000,"publicationDate":"2021-01-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":"{\"title\":\"Transfer learning from speech to music: towards language-sensitive emotion recognition models\",\"authors\":\"Juan Sebastián Gómez Cañón, Estefanía Cano, P. Herrera, E. Gómez\",\"doi\":\"10.23919/Eusipco47968.2020.9287548\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In this study, we address emotion recognition using unsupervised feature learning from speech data, and test its transferability to music. Our approach is to pre-train models using speech in English and Mandarin, and then fine-tune them with excerpts of music labeled with categories of emotion. Our initial hypothesis is that features automatically learned from speech should be transferable to music. Namely, we expect the intra-linguistic setting (e.g., pre-training on speech in English and fine-tuning on music in English) should result in improved performance over the cross-linguistic setting (e.g., pre-training on speech in English and fine-tuning on music in Mandarin). Our results confirm previous research on cross-domain transferability, and encourage research towards language-sensitive Music Emotion Recognition (MER) models.\",\"PeriodicalId\":6705,\"journal\":{\"name\":\"2020 28th European Signal Processing Conference (EUSIPCO)\",\"volume\":\"50 1\",\"pages\":\"136-140\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2021-01-24\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"1\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2020 28th European Signal Processing Conference (EUSIPCO)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.23919/Eusipco47968.2020.9287548\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2020 28th European Signal Processing Conference (EUSIPCO)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.23919/Eusipco47968.2020.9287548","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 1

摘要

在本研究中，我们使用语音数据中的无监督特征学习来解决情感识别问题，并测试其对音乐的可转移性。我们的方法是使用英语和普通话语音对模型进行预训练，然后用标记有情感类别的音乐片段对模型进行微调。我们最初的假设是，从语音中自动学习到的特征应该可以转移到音乐中。也就是说，我们期望语言内设置(例如，英语语音的预训练和英语音乐的微调)应该比跨语言设置(例如，英语语音的预训练和中文音乐的微调)产生更好的性能。我们的研究结果证实了之前关于跨领域可转移性的研究，并鼓励对语言敏感的音乐情感识别(MER)模型的研究。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Transfer learning from speech to music: towards language-sensitive emotion recognition models

In this study, we address emotion recognition using unsupervised feature learning from speech data, and test its transferability to music. Our approach is to pre-train models using speech in English and Mandarin, and then fine-tune them with excerpts of music labeled with categories of emotion. Our initial hypothesis is that features automatically learned from speech should be transferable to music. Namely, we expect the intra-linguistic setting (e.g., pre-training on speech in English and fine-tuning on music in English) should result in improved performance over the cross-linguistic setting (e.g., pre-training on speech in English and fine-tuning on music in Mandarin). Our results confirm previous research on cross-domain transferability, and encourage research towards language-sensitive Music Emotion Recognition (MER) models.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2020 28th European Signal Processing Conference (EUSIPCO)

自引率

0.00%

发文量

期刊最新文献

Eusipco 2021 Cover Page A graph-theoretic sensor-selection scheme for covariance-based Motor Imagery (MI) decoding Hidden Markov Model Based Data-driven Calibration of Non-dispersive Infrared Gas Sensor Deep Transform Learning for Multi-Sensor Fusion Two Stages Parallel LMS Structure: A Pipelined Hardware Architecture