交叉验证策略影响机器学习模型的性能和解释

Lily-belle Sweet, Christoph Müller, Mohit Anand, J. Zscheischler
{"title":"交叉验证策略影响机器学习模型的性能和解释","authors":"Lily-belle Sweet, Christoph Müller, Mohit Anand, J. Zscheischler","doi":"10.1175/aies-d-23-0026.1","DOIUrl":null,"url":null,"abstract":"\nMachine learning algorithms are able to capture complex, nonlinear interacting relationships and are increasingly used to predict yield variability at regional and national scales. Using explainable artificial intelligence (XAI) methods applied to such algorithms may enable better scientific understanding of drivers of yield variability. However, XAI methods may provide misleading results when applied to spatiotemporal correlated datasets. In this study, machine learning models are trained to predict simulated crop yield from climate indices, and the impact of model evaluation strategy on the interpretation and performance of the resulting models is assessed. Using data from a process-based crop model allows us to then comment on the plausibility of the ‘explanations’ provided by XAI methods. Our results show that the choice of evaluation strategy has an impact on (i) interpretations of the model and (ii) model skill on heldout years and regions, after the evaluation strategy is used for hyperparameter-tuning and feature-selection. We find that use of a cross-validation strategy based on clustering in feature-space achieves the most plausible interpretations as well as the best model performance on heldout years and regions. Our results provide first steps towards identifying domain-specific ‘best practices’ for the use of XAI tools on spatiotemporal agricultural or climatic data.","PeriodicalId":94369,"journal":{"name":"Artificial intelligence for the earth systems","volume":"32 1","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2023-07-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":"{\"title\":\"Cross-validation strategy impacts the performance and interpretation of machine learning models\",\"authors\":\"Lily-belle Sweet, Christoph Müller, Mohit Anand, J. Zscheischler\",\"doi\":\"10.1175/aies-d-23-0026.1\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"\\nMachine learning algorithms are able to capture complex, nonlinear interacting relationships and are increasingly used to predict yield variability at regional and national scales. Using explainable artificial intelligence (XAI) methods applied to such algorithms may enable better scientific understanding of drivers of yield variability. However, XAI methods may provide misleading results when applied to spatiotemporal correlated datasets. In this study, machine learning models are trained to predict simulated crop yield from climate indices, and the impact of model evaluation strategy on the interpretation and performance of the resulting models is assessed. Using data from a process-based crop model allows us to then comment on the plausibility of the ‘explanations’ provided by XAI methods. Our results show that the choice of evaluation strategy has an impact on (i) interpretations of the model and (ii) model skill on heldout years and regions, after the evaluation strategy is used for hyperparameter-tuning and feature-selection. We find that use of a cross-validation strategy based on clustering in feature-space achieves the most plausible interpretations as well as the best model performance on heldout years and regions. Our results provide first steps towards identifying domain-specific ‘best practices’ for the use of XAI tools on spatiotemporal agricultural or climatic data.\",\"PeriodicalId\":94369,\"journal\":{\"name\":\"Artificial intelligence for the earth systems\",\"volume\":\"32 1\",\"pages\":\"\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2023-07-10\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"2\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Artificial intelligence for the earth systems\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1175/aies-d-23-0026.1\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Artificial intelligence for the earth systems","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1175/aies-d-23-0026.1","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 2

摘要

机器学习算法能够捕捉复杂的非线性相互作用关系,并越来越多地用于预测区域和国家尺度上的产量变化。将可解释的人工智能(XAI)方法应用于此类算法,可以更好地科学理解产量变化的驱动因素。然而,当应用于时空相关数据集时,XAI方法可能会提供误导性的结果。在本研究中,通过训练机器学习模型来根据气候指数预测模拟作物产量,并评估模型评估策略对结果模型的解释和性能的影响。使用来自基于过程的裁剪模型的数据,我们可以对XAI方法提供的“解释”的合理性进行评论。我们的研究结果表明,在使用评估策略进行超参数调整和特征选择后,评估策略的选择会影响(i)模型的解释和(ii)模型技能对保留年份和地区的影响。我们发现,在特征空间中使用基于聚类的交叉验证策略可以获得最合理的解释,以及在滞留年份和区域上的最佳模型性能。我们的结果为在时空农业或气候数据上使用XAI工具确定特定领域的“最佳实践”提供了第一步。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
Cross-validation strategy impacts the performance and interpretation of machine learning models
Machine learning algorithms are able to capture complex, nonlinear interacting relationships and are increasingly used to predict yield variability at regional and national scales. Using explainable artificial intelligence (XAI) methods applied to such algorithms may enable better scientific understanding of drivers of yield variability. However, XAI methods may provide misleading results when applied to spatiotemporal correlated datasets. In this study, machine learning models are trained to predict simulated crop yield from climate indices, and the impact of model evaluation strategy on the interpretation and performance of the resulting models is assessed. Using data from a process-based crop model allows us to then comment on the plausibility of the ‘explanations’ provided by XAI methods. Our results show that the choice of evaluation strategy has an impact on (i) interpretations of the model and (ii) model skill on heldout years and regions, after the evaluation strategy is used for hyperparameter-tuning and feature-selection. We find that use of a cross-validation strategy based on clustering in feature-space achieves the most plausible interpretations as well as the best model performance on heldout years and regions. Our results provide first steps towards identifying domain-specific ‘best practices’ for the use of XAI tools on spatiotemporal agricultural or climatic data.
求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
自引率
0.00%
发文量
0
期刊最新文献
Transferability and explainability of deep learning emulators for regional climate model projections: Perspectives for future applications Classification of ice particle shapes using machine learning on forward light scattering images Convolutional encoding and normalizing flows: a deep learning approach for offshore wind speed probabilistic forecasting in the Mediterranean Sea Neural networks to find the optimal forcing for offsetting the anthropogenic climate change effects Machine Learning Approach for Spatiotemporal Multivariate Optimization of Environmental Monitoring Sensor Locations
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1