Improving building extraction from high-resolution aerial images: Error correction and performance enhancement using deep learning on the Inria dataset.

IF 2.6 4区综合性期刊 Q2 MULTIDISCIPLINARY SCIENCES Science Progress Pub Date : 2025-01-01 DOI:10.1177/00368504251318202

Serdar Ekiz, Ugur Acar

{"title":"Improving building extraction from high-resolution aerial images: Error correction and performance enhancement using deep learning on the Inria dataset.","authors":"Serdar Ekiz, Ugur Acar","doi":"10.1177/00368504251318202","DOIUrl":null,"url":null,"abstract":"<p><p>Extracting buildings from images is crucial for urban management, urban planning, and post-disaster change detection. Over the years, various approaches have been tried, but the recent application of deep learning has greatly improved the success of such studies. In this study, the Inria dataset was used, consisting of 180 high-resolution aerial images.The study compared the performance of various architectures. DeepLabv3+ emerged as the most successful, with Accuracy, IoU, and F1 Scores of 96.77%, 89.85%, and 94.53%, respectively. Attention U-Net followed, scoring 95.31%, 85.49%, and 91.95%. U-Net, tested with different encoders, achieved average results of 97.22%, 84.78%, and 90.79%. SE-ResNeXt-50 was the best-performing encoder, followed by SE-ResNet-50, ResNeXt-50, and ResNet-50. UNet++ achieved 94.48% Accuracy, 83.09% IoU, and 90.45% F1 Score, while U2Net obtained 94.09%, 82.26%, and 89.88%, making them less successful.When examining the models under challenging conditions, SE-ResNeXt-50 was the most robust, successfully handling scenarios like occlusion by trees and complex indoor gardens. Conversely, Attention U-Net and UNet++ were more prone to errors, particularly when vehicles were parked near buildings or in the presence of shipping containers, where false positives were common. ResNet-50 struggled with concrete gardens, while U2Net showed better results in scenarios involving indoor gardens.These results, compared to other studies using the same dataset with different pixel sizes, show that eliminating erroneous data and resizing images can enhance the performance of deep learning networks. Therefore, by refining the data and adjusting the image sizes, models can make more accurate and efficient building detections.</p>","PeriodicalId":56061,"journal":{"name":"Science Progress","volume":"108 1","pages":"368504251318202"},"PeriodicalIF":2.6000,"publicationDate":"2025-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11822834/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Science Progress","FirstCategoryId":"103","ListUrlMain":"https://doi.org/10.1177/00368504251318202","RegionNum":4,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"MULTIDISCIPLINARY SCIENCES","Score":null,"Total":0}

引用次数: 0

Abstract

Extracting buildings from images is crucial for urban management, urban planning, and post-disaster change detection. Over the years, various approaches have been tried, but the recent application of deep learning has greatly improved the success of such studies. In this study, the Inria dataset was used, consisting of 180 high-resolution aerial images.The study compared the performance of various architectures. DeepLabv3+ emerged as the most successful, with Accuracy, IoU, and F1 Scores of 96.77%, 89.85%, and 94.53%, respectively. Attention U-Net followed, scoring 95.31%, 85.49%, and 91.95%. U-Net, tested with different encoders, achieved average results of 97.22%, 84.78%, and 90.79%. SE-ResNeXt-50 was the best-performing encoder, followed by SE-ResNet-50, ResNeXt-50, and ResNet-50. UNet++ achieved 94.48% Accuracy, 83.09% IoU, and 90.45% F1 Score, while U2Net obtained 94.09%, 82.26%, and 89.88%, making them less successful.When examining the models under challenging conditions, SE-ResNeXt-50 was the most robust, successfully handling scenarios like occlusion by trees and complex indoor gardens. Conversely, Attention U-Net and UNet++ were more prone to errors, particularly when vehicles were parked near buildings or in the presence of shipping containers, where false positives were common. ResNet-50 struggled with concrete gardens, while U2Net showed better results in scenarios involving indoor gardens.These results, compared to other studies using the same dataset with different pixel sizes, show that eliminating erroneous data and resizing images can enhance the performance of deep learning networks. Therefore, by refining the data and adjusting the image sizes, models can make more accurate and efficient building detections.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

求助全文

约1分钟内获得全文去求助

来源期刊

Science Progress Multidisciplinary-Multidisciplinary

CiteScore

3.80

自引率

0.00%

发文量

119

期刊介绍： Science Progress has for over 100 years been a highly regarded review publication in science, technology and medicine. Its objective is to excite the readers'' interest in areas with which they may not be fully familiar but which could facilitate their interest, or even activity, in a cognate field.