{"title":"VITA Search -一个在线媒体资源的智能多模式搜索和存档系统","authors":"Zhanibek Kozhirbayev, Zhandos Yessenbayev, Bagdat Myrzakhmetov","doi":"10.1109/AICT47866.2019.8981781","DOIUrl":null,"url":null,"abstract":"In this paper we present work on intelligent multimodal search and archive system, in which the scientific findings obtained in the work on recognition of Kazakh and Russian speeches, language identification and spoken term detection methods were applied. The paper describes the goals and objectives, the architecture, as well as the subsystem modules of the developed system. The VITA Search system allows for accurately determining the exact time of the required spoken information in the data in Kazakh and Russian languages from various broadcast channels. The speech recognition unit uses the Kaldi toolkit to generate lattices from the raw audio data. An acoustic model trained using deep neural networks shows significant results. The word error rate on the train set for recognition of Kazakh speech was 3.86, and for Russian speech - 9.85. Moreover, we integrated a language identification model trained using Long Short-Term Memory Recurrent Neural Networks in order to select the correct model for the input audio. Regarding spoken term detection, we applied word and proxy-based approaches to search for keyword terms among the lattices.","PeriodicalId":329473,"journal":{"name":"2019 IEEE 13th International Conference on Application of Information and Communication Technologies (AICT)","volume":"279 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2019-10-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"VITA Search - An Intelligent Multimodal Search and Archive System for Online Media Resources\",\"authors\":\"Zhanibek Kozhirbayev, Zhandos Yessenbayev, Bagdat Myrzakhmetov\",\"doi\":\"10.1109/AICT47866.2019.8981781\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In this paper we present work on intelligent multimodal search and archive system, in which the scientific findings obtained in the work on recognition of Kazakh and Russian speeches, language identification and spoken term detection methods were applied. The paper describes the goals and objectives, the architecture, as well as the subsystem modules of the developed system. The VITA Search system allows for accurately determining the exact time of the required spoken information in the data in Kazakh and Russian languages from various broadcast channels. The speech recognition unit uses the Kaldi toolkit to generate lattices from the raw audio data. An acoustic model trained using deep neural networks shows significant results. The word error rate on the train set for recognition of Kazakh speech was 3.86, and for Russian speech - 9.85. Moreover, we integrated a language identification model trained using Long Short-Term Memory Recurrent Neural Networks in order to select the correct model for the input audio. Regarding spoken term detection, we applied word and proxy-based approaches to search for keyword terms among the lattices.\",\"PeriodicalId\":329473,\"journal\":{\"name\":\"2019 IEEE 13th International Conference on Application of Information and Communication Technologies (AICT)\",\"volume\":\"279 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2019-10-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2019 IEEE 13th International Conference on Application of Information and Communication Technologies (AICT)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/AICT47866.2019.8981781\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2019 IEEE 13th International Conference on Application of Information and Communication Technologies (AICT)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/AICT47866.2019.8981781","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
VITA Search - An Intelligent Multimodal Search and Archive System for Online Media Resources
In this paper we present work on intelligent multimodal search and archive system, in which the scientific findings obtained in the work on recognition of Kazakh and Russian speeches, language identification and spoken term detection methods were applied. The paper describes the goals and objectives, the architecture, as well as the subsystem modules of the developed system. The VITA Search system allows for accurately determining the exact time of the required spoken information in the data in Kazakh and Russian languages from various broadcast channels. The speech recognition unit uses the Kaldi toolkit to generate lattices from the raw audio data. An acoustic model trained using deep neural networks shows significant results. The word error rate on the train set for recognition of Kazakh speech was 3.86, and for Russian speech - 9.85. Moreover, we integrated a language identification model trained using Long Short-Term Memory Recurrent Neural Networks in order to select the correct model for the input audio. Regarding spoken term detection, we applied word and proxy-based approaches to search for keyword terms among the lattices.