{"title":"Application and Research of Music Generation System Based on CVAE and Transformer-XL in Video Background Music","authors":"Jun Min;Zhiwei Gao;Lei Wang","doi":"10.1109/TII.2024.3477561","DOIUrl":null,"url":null,"abstract":"In the field of music generation using algorithms, processing time-series data has consistently been a complex task. To improve music generation with long sequences, insightful-unit-conditional variational autoencoder is proposed, which can enhance unit-conditional variational autoencoders with an improved attention mechanism. This model integrates TransformerXLs recurrent mechanism and relative positional encoding with measure-level granularity. For practical applications, a scheme is addressed that uses optical flow to extract motion features from video frames, quantifying motion rate and intensity. Furthermore, a dynamic correlation method is proposed to align video motion features with musical rhythm, guiding the model to generate melodies that match the videos rhythm.","PeriodicalId":13301,"journal":{"name":"IEEE Transactions on Industrial Informatics","volume":"21 2","pages":"1409-1418"},"PeriodicalIF":9.9000,"publicationDate":"2024-10-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"IEEE Transactions on Industrial Informatics","FirstCategoryId":"94","ListUrlMain":"https://ieeexplore.ieee.org/document/10729273/","RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"AUTOMATION & CONTROL SYSTEMS","Score":null,"Total":0}
引用次数: 0
Abstract
In the field of music generation using algorithms, processing time-series data has consistently been a complex task. To improve music generation with long sequences, insightful-unit-conditional variational autoencoder is proposed, which can enhance unit-conditional variational autoencoders with an improved attention mechanism. This model integrates TransformerXLs recurrent mechanism and relative positional encoding with measure-level granularity. For practical applications, a scheme is addressed that uses optical flow to extract motion features from video frames, quantifying motion rate and intensity. Furthermore, a dynamic correlation method is proposed to align video motion features with musical rhythm, guiding the model to generate melodies that match the videos rhythm.
期刊介绍:
The IEEE Transactions on Industrial Informatics is a multidisciplinary journal dedicated to publishing technical papers that connect theory with practical applications of informatics in industrial settings. It focuses on the utilization of information in intelligent, distributed, and agile industrial automation and control systems. The scope includes topics such as knowledge-based and AI-enhanced automation, intelligent computer control systems, flexible and collaborative manufacturing, industrial informatics in software-defined vehicles and robotics, computer vision, industrial cyber-physical and industrial IoT systems, real-time and networked embedded systems, security in industrial processes, industrial communications, systems interoperability, and human-machine interaction.