An outlier detection framework for Air Quality Index prediction using linear and ensemble models

Pradeep Kumar Dongre , Viral Patel , Upendra Bhoi , Nilesh N. Maltare
{"title":"An outlier detection framework for Air Quality Index prediction using linear and ensemble models","authors":"Pradeep Kumar Dongre ,&nbsp;Viral Patel ,&nbsp;Upendra Bhoi ,&nbsp;Nilesh N. Maltare","doi":"10.1016/j.dajour.2025.100546","DOIUrl":null,"url":null,"abstract":"<div><div>The Air Quality Index (AQI) is a key indicator for assessing air quality and its associated health impacts. Accurate AQI calculations are crucial for reliable air quality assessments, but outliers in air quality data can distort these calculations, leading to inaccurate predictions. This paper presents a comprehensive framework for air quality prediction that integrates multiple outlier detection methods with machine learning models, focusing on enhancing the accuracy and robustness of predictions. The study investigates various outlier detection techniques, including the Interquartile Range (IQR), robust Z-score, and Mahalanobis distance, and evaluates their impact when integrated into machine learning models. Unlike traditional approaches that remove outliers without considering seasonal effects, this research proposes retaining extreme data points after seasonal validation to improve model generalization and prediction accuracy for unseen data. The framework is evaluated using a dataset from Jaipur city, testing multiple machine learning models, including linear regression, ensemble methods, and K-Nearest Neighbor (KNN) regression. Results show that the integrated framework significantly improves model performance, with the Extra Trees Regressor achieving the best results (MAE = 11.9161, RMSE = 16.1660, and <span><math><msup><mrow><mi>R</mi></mrow><mrow><mn>2</mn></mrow></msup></math></span> = 0.8884) after refinement, compared to baseline performance (MAE = 12.6765, RMSE = 17.8452, and <span><math><msup><mrow><mi>R</mi></mrow><mrow><mn>2</mn></mrow></msup></math></span> = 0.8737). This study demonstrates the empirical effectiveness of the proposed framework and provides practical guidelines for air quality prediction in real-world applications.</div></div>","PeriodicalId":100357,"journal":{"name":"Decision Analytics Journal","volume":"14 ","pages":"Article 100546"},"PeriodicalIF":0.0000,"publicationDate":"2025-01-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Decision Analytics Journal","FirstCategoryId":"1085","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S2772662225000025","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0

Abstract

The Air Quality Index (AQI) is a key indicator for assessing air quality and its associated health impacts. Accurate AQI calculations are crucial for reliable air quality assessments, but outliers in air quality data can distort these calculations, leading to inaccurate predictions. This paper presents a comprehensive framework for air quality prediction that integrates multiple outlier detection methods with machine learning models, focusing on enhancing the accuracy and robustness of predictions. The study investigates various outlier detection techniques, including the Interquartile Range (IQR), robust Z-score, and Mahalanobis distance, and evaluates their impact when integrated into machine learning models. Unlike traditional approaches that remove outliers without considering seasonal effects, this research proposes retaining extreme data points after seasonal validation to improve model generalization and prediction accuracy for unseen data. The framework is evaluated using a dataset from Jaipur city, testing multiple machine learning models, including linear regression, ensemble methods, and K-Nearest Neighbor (KNN) regression. Results show that the integrated framework significantly improves model performance, with the Extra Trees Regressor achieving the best results (MAE = 11.9161, RMSE = 16.1660, and R2 = 0.8884) after refinement, compared to baseline performance (MAE = 12.6765, RMSE = 17.8452, and R2 = 0.8737). This study demonstrates the empirical effectiveness of the proposed framework and provides practical guidelines for air quality prediction in real-world applications.
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
求助全文
约1分钟内获得全文 去求助
来源期刊
CiteScore
3.90
自引率
0.00%
发文量
0
期刊最新文献
A comprehensive review of dwarf mongoose optimization algorithm with emerging trends and future research directions A hybrid multi-objective optimization approach with NSGA-II for feature selection A novel Full Multiplicative Data Envelopment Analysis Model for solving Multi-Attribute Decision-Making problems An investigation of supervised machine learning models for predicting drivers’ ethical decisions in autonomous vehicles An outlier detection framework for Air Quality Index prediction using linear and ensemble models
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1