{"title":"Common issues of data science on the eco-environmental risks of emerging contaminants","authors":"Xiangang Hu, Xu Dong, Zhangjia Wang","doi":"10.1016/j.envint.2025.109301","DOIUrl":null,"url":null,"abstract":"<div><div>Data-driven approaches (e.g., machine learning) are increasingly used to replace or assist laboratory studies in the study of emerging contaminants (ECs). In the past ten years, an increasing number of models or approaches have been applied to ECs, and the datasets used are continuously enriched. However, there are large knowledge gaps between what we have found and the natural eco-environmental meaning. For most published reviews, the contents are organized by the types of ECs, but the common issues of data science, regardless of the type of pollutant, are not sufficiently addressed. To close or narrow the knowledge gaps, we highlight the following issues ignored in the field of data-driven EC research. Complicated biological and ecological data and ensemble models revealing mechanisms and spatiotemporal trends with strong causal relationships and without data leakage deserve more attention in the future. In addition, the matrix influence, trace concentration, and complex scenario have often been ignored in previous works. Therefore, an integrated research framework related to natural fields, ecological systems, and large-scale environmental problems, rather than relying solely on laboratory data-related analysis, is urgently needed. Beyond the current prediction purposes, data science can inspire the discovery of scientific questions, and mutual inspiration among data science, process and mechanism models, and laboratory and field research is a critical direction. Focusing on the above urgent and common issues related to data, frameworks, and purposes, regardless of the type of pollutant, data science is expected to achieve great advancements in addressing the eco-environmental risks of ECs.</div></div>","PeriodicalId":308,"journal":{"name":"Environment International","volume":"196 ","pages":"Article 109301"},"PeriodicalIF":10.3000,"publicationDate":"2025-02-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Environment International","FirstCategoryId":"93","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S0160412025000522","RegionNum":1,"RegionCategory":"环境科学与生态学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"ENVIRONMENTAL SCIENCES","Score":null,"Total":0}
引用次数: 0
Abstract
Data-driven approaches (e.g., machine learning) are increasingly used to replace or assist laboratory studies in the study of emerging contaminants (ECs). In the past ten years, an increasing number of models or approaches have been applied to ECs, and the datasets used are continuously enriched. However, there are large knowledge gaps between what we have found and the natural eco-environmental meaning. For most published reviews, the contents are organized by the types of ECs, but the common issues of data science, regardless of the type of pollutant, are not sufficiently addressed. To close or narrow the knowledge gaps, we highlight the following issues ignored in the field of data-driven EC research. Complicated biological and ecological data and ensemble models revealing mechanisms and spatiotemporal trends with strong causal relationships and without data leakage deserve more attention in the future. In addition, the matrix influence, trace concentration, and complex scenario have often been ignored in previous works. Therefore, an integrated research framework related to natural fields, ecological systems, and large-scale environmental problems, rather than relying solely on laboratory data-related analysis, is urgently needed. Beyond the current prediction purposes, data science can inspire the discovery of scientific questions, and mutual inspiration among data science, process and mechanism models, and laboratory and field research is a critical direction. Focusing on the above urgent and common issues related to data, frameworks, and purposes, regardless of the type of pollutant, data science is expected to achieve great advancements in addressing the eco-environmental risks of ECs.
期刊介绍:
Environmental Health publishes manuscripts focusing on critical aspects of environmental and occupational medicine, including studies in toxicology and epidemiology, to illuminate the human health implications of exposure to environmental hazards. The journal adopts an open-access model and practices open peer review.
It caters to scientists and practitioners across all environmental science domains, directly or indirectly impacting human health and well-being. With a commitment to enhancing the prevention of environmentally-related health risks, Environmental Health serves as a public health journal for the community and scientists engaged in matters of public health significance concerning the environment.