{"title":"QPWS Feature Selection and CAE Fusion of Visible/Near-Infrared Spectroscopy Data for the Identification of Salix psammophila Origin","authors":"Yicheng Ma, Ying Li, Xinkai Peng, Congyu Chen, Hengkai Li, Xinping Wang, Weilong Wang, Xiaozhen Lan, Jixuan Wang, Zhiyong Pei","doi":"10.3390/f15010006","DOIUrl":null,"url":null,"abstract":"Salix psammophila, classified under the Salicaceae family, is a deciduous, densely branched, and erect shrub. As a leading pioneer tree species in windbreak and sand stabilization, it has played a crucial role in combating desertification in northwestern China. However, different genetic sources of Salix psammophila exhibit significant variations in their effectiveness for windbreak and sand stabilization. Therefore, it is essential to establish a rapid and reliable method for identifying different Salix psammophila varieties. Visible and near-infrared (Vis-NIR) spectroscopy is currently a reliable non-destructive solution for origin traceability. This study introduced a novel feature selection strategy, called qualitative percentile weighted sampling (QPWS), based on the principle of the long tail effect for Vis-NIR spectroscopy. The core idea of QPWS combines weighted sampling and percentage wavelength selection to identify key wavelengths. By employing a multi-threaded parallel execution of multiple QPWS instances, we aimed to search for the optimal feature bands to address the instability issues that can arise during the feature selection process. To address the problem of reduced prediction performance in one-dimensional convolutional neural network (1D-CNN) models after feature selection, we have introduced convolutional autoencoders (CAEs) to reduce the dimensions of wavelengths that are discarded during feature selection. Subsequently, these reduced dimensions are fused with the selected wavelengths, thereby enhancing the model’s performance. With our completed model, we selected outstanding models for model fusion and established a decision system for Salix psammophila. It is worth noting that all 1D-CNN models in this study were developed using Bayesian optimization methods. In comparison with principal component analysis (PCA) and full spectrum methods, QPWS exhibits superior predictive performance in the field of machine learning. In the realm of deep learning, the fusion of data combining QPWS with CAE demonstrated even greater potential with an improvement of average accuracy of approximately 2.13% when compared to QPWS alone and a 228% increase in operational speed compared to a model with full spectra. These results indicated that the combination of CAE with QPWS can be an effective tool for identifying the origin of Salix psammophila.","PeriodicalId":12339,"journal":{"name":"Forests","volume":" 2","pages":""},"PeriodicalIF":2.4000,"publicationDate":"2023-12-19","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Forests","FirstCategoryId":"97","ListUrlMain":"https://doi.org/10.3390/f15010006","RegionNum":2,"RegionCategory":"农林科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"FORESTRY","Score":null,"Total":0}
引用次数: 0
Abstract
Salix psammophila, classified under the Salicaceae family, is a deciduous, densely branched, and erect shrub. As a leading pioneer tree species in windbreak and sand stabilization, it has played a crucial role in combating desertification in northwestern China. However, different genetic sources of Salix psammophila exhibit significant variations in their effectiveness for windbreak and sand stabilization. Therefore, it is essential to establish a rapid and reliable method for identifying different Salix psammophila varieties. Visible and near-infrared (Vis-NIR) spectroscopy is currently a reliable non-destructive solution for origin traceability. This study introduced a novel feature selection strategy, called qualitative percentile weighted sampling (QPWS), based on the principle of the long tail effect for Vis-NIR spectroscopy. The core idea of QPWS combines weighted sampling and percentage wavelength selection to identify key wavelengths. By employing a multi-threaded parallel execution of multiple QPWS instances, we aimed to search for the optimal feature bands to address the instability issues that can arise during the feature selection process. To address the problem of reduced prediction performance in one-dimensional convolutional neural network (1D-CNN) models after feature selection, we have introduced convolutional autoencoders (CAEs) to reduce the dimensions of wavelengths that are discarded during feature selection. Subsequently, these reduced dimensions are fused with the selected wavelengths, thereby enhancing the model’s performance. With our completed model, we selected outstanding models for model fusion and established a decision system for Salix psammophila. It is worth noting that all 1D-CNN models in this study were developed using Bayesian optimization methods. In comparison with principal component analysis (PCA) and full spectrum methods, QPWS exhibits superior predictive performance in the field of machine learning. In the realm of deep learning, the fusion of data combining QPWS with CAE demonstrated even greater potential with an improvement of average accuracy of approximately 2.13% when compared to QPWS alone and a 228% increase in operational speed compared to a model with full spectra. These results indicated that the combination of CAE with QPWS can be an effective tool for identifying the origin of Salix psammophila.
期刊介绍:
Forests (ISSN 1999-4907) is an international and cross-disciplinary scholarly journal of forestry and forest ecology. It publishes research papers, short communications and review papers. There is no restriction on the length of the papers. Our aim is to encourage scientists to publish their experimental and theoretical research in as much detail as possible. Full experimental and/or methodical details must be provided for research articles.