Pub Date : 2025-10-30DOI: 10.1109/LGRS.2025.3626786
Binge Cui;Shengyun Liu;Jing Zhang;Yan Lu
Coastline extraction from remote sensing imagery is persistently challenged by intra-class heterogeneity (e.g., diverse coastline types) and boundary ambiguity. Existing methods often exhibit suboptimal performance in complex scenes mixing artificial and natural landforms, as they tend to ignore coastline morphological priors and struggle to recover details in low-contrast regions. To address these issues, this letter introduces TopoSegNet, a novel collaborative framework centered on a dual-decoder architecture. A segmentation decoder utilizes a morphology-aware attention (MAA) module to adaptively decouple and model diverse coastline morphologies and a structure-detail synergistic enhancement (SDSE) module to reconstruct weak boundaries with high fidelity. Meanwhile, a learnable topology decoder frames topology construction as a graph reasoning task, which ensures the geometric and topological integrity of the final vector output. TopoSegNet was evaluated on the public Landsat-8 and a custom Lianyungang Gaofen-1 (GF-1) dataset. The experimental results show that the proposed method reached 98.64%, 66.80%, and 0.795% on the mIoU, BIoU, and average path length similarity (APLS) metrics, respectively, verifying its validity and superiority. Compared to the state-of-the-art methods, the TopoSegNet model demonstrates significantly higher accuracy and topological fidelity.
{"title":"TopoSegNet: Enhancing Geometric Fidelity of Coastline Extraction via a Joint Segmentation and Topological Reasoning Framework","authors":"Binge Cui;Shengyun Liu;Jing Zhang;Yan Lu","doi":"10.1109/LGRS.2025.3626786","DOIUrl":"https://doi.org/10.1109/LGRS.2025.3626786","url":null,"abstract":"Coastline extraction from remote sensing imagery is persistently challenged by intra-class heterogeneity (e.g., diverse coastline types) and boundary ambiguity. Existing methods often exhibit suboptimal performance in complex scenes mixing artificial and natural landforms, as they tend to ignore coastline morphological priors and struggle to recover details in low-contrast regions. To address these issues, this letter introduces TopoSegNet, a novel collaborative framework centered on a dual-decoder architecture. A segmentation decoder utilizes a morphology-aware attention (MAA) module to adaptively decouple and model diverse coastline morphologies and a structure-detail synergistic enhancement (SDSE) module to reconstruct weak boundaries with high fidelity. Meanwhile, a learnable topology decoder frames topology construction as a graph reasoning task, which ensures the geometric and topological integrity of the final vector output. TopoSegNet was evaluated on the public Landsat-8 and a custom Lianyungang Gaofen-1 (GF-1) dataset. The experimental results show that the proposed method reached 98.64%, 66.80%, and 0.795% on the mIoU, BIoU, and average path length similarity (APLS) metrics, respectively, verifying its validity and superiority. Compared to the state-of-the-art methods, the TopoSegNet model demonstrates significantly higher accuracy and topological fidelity.","PeriodicalId":91017,"journal":{"name":"IEEE geoscience and remote sensing letters : a publication of the IEEE Geoscience and Remote Sensing Society","volume":"23 ","pages":"1-5"},"PeriodicalIF":4.4,"publicationDate":"2025-10-30","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145612124","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Deep learning approaches that jointly learn feature extraction have achieved remarkable progress in image matching. However, current methods often treat central and neighboring pixels uniformly and use static feature selection strategies that fail to account for environmental variations. This results in limited robustness of descriptors and keypoints, thereby affecting matching accuracy. To address these limitations, we propose a robust joint optimization network for feature detection and description in optical and SAR image matching. A center-weighted module (CWM) is designed to enhance local feature representation by emphasizing the hierarchical relationship between central and surrounding features. Furthermore, a multiscale gated aggregation (MSGA) module is introduced to suppress redundant responses and improve keypoint discriminability through a gating mechanism. To address the inconsistency of score maps across heterogeneous modalities, we design a position-constrained repeatability loss to guide the network in learning stable and consistent keypoint correspondences. Experimental results across various scenarios demonstrate that the proposed method outperforms state-of-the-art techniques in terms of both matching accuracy and the number of correct matches, highlighting its robustness and effectiveness.
{"title":"A Robust Joint Optimization Network for Feature Detection and Description in Optical and SAR Image Matching","authors":"Xinshan Zhang;Zhitao Fu;Menghua Li;Shaochen Zhang;Han Nie;Bo-Hui Tang","doi":"10.1109/LGRS.2025.3626750","DOIUrl":"https://doi.org/10.1109/LGRS.2025.3626750","url":null,"abstract":"Deep learning approaches that jointly learn feature extraction have achieved remarkable progress in image matching. However, current methods often treat central and neighboring pixels uniformly and use static feature selection strategies that fail to account for environmental variations. This results in limited robustness of descriptors and keypoints, thereby affecting matching accuracy. To address these limitations, we propose a robust joint optimization network for feature detection and description in optical and SAR image matching. A center-weighted module (CWM) is designed to enhance local feature representation by emphasizing the hierarchical relationship between central and surrounding features. Furthermore, a multiscale gated aggregation (MSGA) module is introduced to suppress redundant responses and improve keypoint discriminability through a gating mechanism. To address the inconsistency of score maps across heterogeneous modalities, we design a position-constrained repeatability loss to guide the network in learning stable and consistent keypoint correspondences. Experimental results across various scenarios demonstrate that the proposed method outperforms state-of-the-art techniques in terms of both matching accuracy and the number of correct matches, highlighting its robustness and effectiveness.","PeriodicalId":91017,"journal":{"name":"IEEE geoscience and remote sensing letters : a publication of the IEEE Geoscience and Remote Sensing Society","volume":"23 ","pages":"1-5"},"PeriodicalIF":4.4,"publicationDate":"2025-10-30","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145537627","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2025-10-28DOI: 10.1109/LGRS.2025.3626369
Haowen Jin;Yuankang Ye;Chang Liu;Feng Gao
Precipitation nowcasting using radar echo data is critical for issuing timely extreme weather warnings, yet the existing models struggle to balance computational efficiency with prediction accuracy when modeling complex, nonlinear echo sequences. To address these challenges, we propose MambaCast, a novel dual-branch precipitation nowcasting model built upon the Mamba framework. Specifically, MambaCast incorporates three key components: a state-space model (SSM) branch, a convolutional neural network (CNN) branch and a CastFusion module. The SSM branch captures global low-frequency evolution features in the radar echo field through a selective scanning mechanism, while the CNN branch extracts local high-frequency transient features using gated spatiotemporal attention (gSTA). The CastFusion module dynamically integrates features across different frequency scales, enabling adaptive fusion of spatiotemporal distribution. Experiments on two public radar datasets show that MambaCast consistently outperforms baseline models.
{"title":"MambaCast: An Efficient Precipitation Nowcasting Model With Dual-Branch Mamba","authors":"Haowen Jin;Yuankang Ye;Chang Liu;Feng Gao","doi":"10.1109/LGRS.2025.3626369","DOIUrl":"https://doi.org/10.1109/LGRS.2025.3626369","url":null,"abstract":"Precipitation nowcasting using radar echo data is critical for issuing timely extreme weather warnings, yet the existing models struggle to balance computational efficiency with prediction accuracy when modeling complex, nonlinear echo sequences. To address these challenges, we propose MambaCast, a novel dual-branch precipitation nowcasting model built upon the Mamba framework. Specifically, MambaCast incorporates three key components: a state-space model (SSM) branch, a convolutional neural network (CNN) branch and a CastFusion module. The SSM branch captures global low-frequency evolution features in the radar echo field through a selective scanning mechanism, while the CNN branch extracts local high-frequency transient features using gated spatiotemporal attention (gSTA). The CastFusion module dynamically integrates features across different frequency scales, enabling adaptive fusion of spatiotemporal distribution. Experiments on two public radar datasets show that MambaCast consistently outperforms baseline models.","PeriodicalId":91017,"journal":{"name":"IEEE geoscience and remote sensing letters : a publication of the IEEE Geoscience and Remote Sensing Society","volume":"23 ","pages":"1-5"},"PeriodicalIF":4.4,"publicationDate":"2025-10-28","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145537628","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2025-10-20DOI: 10.1109/LGRS.2025.3623244
Aybora Köksal;A. Aydın Alatan
Remote sensing (RS) applications often rely on edge hardware that cannot host the models in the 7B parametric vision language of today. This letter presents TinyRS, the first 2B-parameter vision language models (VLMs) optimized for RS, and TinyRS-R1, its reasoning-augmented variant. Based on Qwen2-VL-2B, TinyRS is trained via a four-stage pipeline: pretraining on million-scale satellite images, instruction tuning, fine-tuning with chain-of-thought (CoT) annotations from a new reasoning dataset, and group relative policy optimization (GRPO)-based alignment. TinyRS-R1 matches or surpasses recent 7B RS models in classification, visual question answering (VQA), grounding, and open-ended QA—while using one third of the memory and latency. CoT reasoning improves grounding and scene understanding, while TinyRS excels at concise, low-latency VQA. TinyRS-R1 is the first domain-specialized small VLM with GRPO-aligned CoT reasoning for general-purpose RS. The code, models, and caption datasets are available at https://github.com/aybora/TinyRS
{"title":"TinyRS-R1: Compact Vision Language Model for Remote Sensing","authors":"Aybora Köksal;A. Aydın Alatan","doi":"10.1109/LGRS.2025.3623244","DOIUrl":"https://doi.org/10.1109/LGRS.2025.3623244","url":null,"abstract":"Remote sensing (RS) applications often rely on edge hardware that cannot host the models in the 7B parametric vision language of today. This letter presents TinyRS, the first 2B-parameter vision language models (VLMs) optimized for RS, and TinyRS-R1, its reasoning-augmented variant. Based on Qwen2-VL-2B, TinyRS is trained via a four-stage pipeline: pretraining on million-scale satellite images, instruction tuning, fine-tuning with chain-of-thought (CoT) annotations from a new reasoning dataset, and group relative policy optimization (GRPO)-based alignment. TinyRS-R1 matches or surpasses recent 7B RS models in classification, visual question answering (VQA), grounding, and open-ended QA—while using one third of the memory and latency. CoT reasoning improves grounding and scene understanding, while TinyRS excels at concise, low-latency VQA. TinyRS-R1 is the first domain-specialized small VLM with GRPO-aligned CoT reasoning for general-purpose RS. The code, models, and caption datasets are available at <uri>https://github.com/aybora/TinyRS</uri>","PeriodicalId":91017,"journal":{"name":"IEEE geoscience and remote sensing letters : a publication of the IEEE Geoscience and Remote Sensing Society","volume":"22 ","pages":"1-5"},"PeriodicalIF":4.4,"publicationDate":"2025-10-20","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145405232","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2025-10-13DOI: 10.1109/LGRS.2025.3620872
Jingfan Wang;Wen Lu;Zeming Zhang;Zhaoyang Wang;Zhe Li
Transformer-based methods for remote sensing image super-resolution (SR) face challenges in reconstructing high-frequency textures due to the interference from large flat regions, such as farmlands and water bodies. To address these limitations, we propose a channel-enhanced multiscale window attention mechanism, which is designed to minimize the impact of flat regions on high-frequency area reconstruction while effectively utilizing the intrinsic multiscale features of remote sensing images. To better capture the multiscale features of remote sensing images, we introduce a series of depthwise separable convolution kernels of varying sizes during the shallow feature extraction stage. Experimental results demonstrate that the proposed method achieves superior peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) scores across multiple remote sensing benchmark datasets and scaling factors, validating its effectiveness.
{"title":"Multiscale Window Attention Channel Enhanced for Remote Sensing Image Super-Resolution","authors":"Jingfan Wang;Wen Lu;Zeming Zhang;Zhaoyang Wang;Zhe Li","doi":"10.1109/LGRS.2025.3620872","DOIUrl":"https://doi.org/10.1109/LGRS.2025.3620872","url":null,"abstract":"Transformer-based methods for remote sensing image super-resolution (SR) face challenges in reconstructing high-frequency textures due to the interference from large flat regions, such as farmlands and water bodies. To address these limitations, we propose a channel-enhanced multiscale window attention mechanism, which is designed to minimize the impact of flat regions on high-frequency area reconstruction while effectively utilizing the intrinsic multiscale features of remote sensing images. To better capture the multiscale features of remote sensing images, we introduce a series of depthwise separable convolution kernels of varying sizes during the shallow feature extraction stage. Experimental results demonstrate that the proposed method achieves superior peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) scores across multiple remote sensing benchmark datasets and scaling factors, validating its effectiveness.","PeriodicalId":91017,"journal":{"name":"IEEE geoscience and remote sensing letters : a publication of the IEEE Geoscience and Remote Sensing Society","volume":"23 ","pages":"1-5"},"PeriodicalIF":4.4,"publicationDate":"2025-10-13","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145778330","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pretrained vision–language models (VLMs) have demonstrated promising performance in remote sensing (RS) image–text retrieval tasks. However, the scarcity of high-quality image–text datasets remains a challenge in fine-tuning VLMs for RS. The captions in existing datasets tend to be uniform and lack details. To fully use rich detailed information from RS images, we propose a method to fine-tune VLMs. We first construct a new visual–language dataset that balances both global and local information for RS (GLRS) image–text retrieval. Specifically, a multimodal large language model (MLLM) is used to generate captions for local patches and global captions for the entire image. To effectively use local information, we propose a global and local image captioning method (GLCap). With a large language model (LLM), we further obtain higher quality captions by merging both global and local captions. Finally, we fine-tune the weights of RS-M-contrastive language image pretraining (CLIP) with a progressive global–local fine-tuning strategy on GLRS. Experimental results demonstrate that our method outperforms state-of-the-art (SoTA) approaches on two common RS image–text retrieval downstream tasks. Our code and dataset are available at https://github.com/hhu-czy/GLRS
预训练的视觉语言模型(VLMs)在遥感图像文本检索任务中表现出了良好的性能。然而,高质量的图像-文本数据集的缺乏仍然是对遥感vlm进行微调的一个挑战,现有数据集的标题往往是统一的,缺乏细节。为了充分利用RS图像中丰富的细节信息,我们提出了一种微调VLMs的方法。我们首先构建了一个新的视觉语言数据集,该数据集平衡了RS (GLRS)图像文本检索的全局和局部信息。具体而言,使用多模态大语言模型(multimodal large language model, MLLM)生成局部补丁的标题和整个图像的全局标题。为了有效地利用局部信息,我们提出了一种全局和局部图像字幕方法(GLCap)。使用大型语言模型(LLM),我们通过合并全局和局部字幕进一步获得更高质量的字幕。最后,我们在GLRS上采用渐进的全局-局部微调策略对rs - m对比语言图像预训练(CLIP)的权重进行微调。实验结果表明,我们的方法在两个常见的RS图像文本检索下游任务上优于最先进的(SoTA)方法。我们的代码和数据集可在https://github.com/hhu-czy/GLRS上获得
{"title":"Integrating Global and Local Information for Remote Sensing Image–Text Retrieval","authors":"Ziyun Chen;Fan Liu;Zhangqingyun Guan;Qian Zhou;Xiaocong Zhou;Chuanyi Zhang","doi":"10.1109/LGRS.2025.3616154","DOIUrl":"https://doi.org/10.1109/LGRS.2025.3616154","url":null,"abstract":"Pretrained vision–language models (VLMs) have demonstrated promising performance in remote sensing (RS) image–text retrieval tasks. However, the scarcity of high-quality image–text datasets remains a challenge in fine-tuning VLMs for RS. The captions in existing datasets tend to be uniform and lack details. To fully use rich detailed information from RS images, we propose a method to fine-tune VLMs. We first construct a new visual–language dataset that balances both global and local information for RS (GLRS) image–text retrieval. Specifically, a multimodal large language model (MLLM) is used to generate captions for local patches and global captions for the entire image. To effectively use local information, we propose a global and local image captioning method (GLCap). With a large language model (LLM), we further obtain higher quality captions by merging both global and local captions. Finally, we fine-tune the weights of RS-M-contrastive language image pretraining (CLIP) with a progressive global–local fine-tuning strategy on GLRS. Experimental results demonstrate that our method outperforms state-of-the-art (SoTA) approaches on two common RS image–text retrieval downstream tasks. Our code and dataset are available at <uri>https://github.com/hhu-czy/GLRS</uri>","PeriodicalId":91017,"journal":{"name":"IEEE geoscience and remote sensing letters : a publication of the IEEE Geoscience and Remote Sensing Society","volume":"22 ","pages":"1-5"},"PeriodicalIF":4.4,"publicationDate":"2025-10-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145455801","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2025-09-12DOI: 10.1109/LGRS.2025.3609444
Li Liu;Yongcheng Zhou;Hang Xu;Jingxia Li;Jianguo Zhang;Lijun Zhou;Bingjie Wang
Automatic underground object classification based on deep learning (DL) has been widely used in ground penetrating radar (GPR) fields. However, its excellent performance heavily depends on sufficient labeled training data. In GPR fields, large amounts of labeled data are difficult to obtain due to time-consuming and experience-dependent manual annotation work. To address the issue of limited labeled data, we propose a novel semi-supervised learning (SSL) method for urban-road underground multiclass object classification. It fully utilizes abundant unlabeled data and limited labeled data to enhance classification performance. We applied a variant of the triple-GAN (TGAN) model and modified it by introducing a similarity constraint, which is associated with GPR data geometric features and can help to produce high-quality generated images. Experimental results of laboratory and field data show that it has higher accuracy than representative baseline methods under limited labeled data.
{"title":"Semi-Supervised Triple-GAN With Similarity Constraint for Automatic Underground Object Classification Using Ground Penetrating Radar Data","authors":"Li Liu;Yongcheng Zhou;Hang Xu;Jingxia Li;Jianguo Zhang;Lijun Zhou;Bingjie Wang","doi":"10.1109/LGRS.2025.3609444","DOIUrl":"https://doi.org/10.1109/LGRS.2025.3609444","url":null,"abstract":"Automatic underground object classification based on deep learning (DL) has been widely used in ground penetrating radar (GPR) fields. However, its excellent performance heavily depends on sufficient labeled training data. In GPR fields, large amounts of labeled data are difficult to obtain due to time-consuming and experience-dependent manual annotation work. To address the issue of limited labeled data, we propose a novel semi-supervised learning (SSL) method for urban-road underground multiclass object classification. It fully utilizes abundant unlabeled data and limited labeled data to enhance classification performance. We applied a variant of the triple-GAN (TGAN) model and modified it by introducing a similarity constraint, which is associated with GPR data geometric features and can help to produce high-quality generated images. Experimental results of laboratory and field data show that it has higher accuracy than representative baseline methods under limited labeled data.","PeriodicalId":91017,"journal":{"name":"IEEE geoscience and remote sensing letters : a publication of the IEEE Geoscience and Remote Sensing Society","volume":"22 ","pages":"1-5"},"PeriodicalIF":4.4,"publicationDate":"2025-09-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145078645","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2025-09-11DOI: 10.1109/LGRS.2025.3608704
Long Tang;Hong Zhang;Yumei Li;Fan Xu;Fang Zou
This study investigates the large-scale ionospheric traveling disturbances (LSTIDs) over North America and Europe associated with the intense geomagnetic storm in May 2024, utilizing total electron content (TEC) data derived from ground-based Global Navigation Satellite System (GNSS) stations. The findings reveal that the observed LSTIDs in both regions exhibited an unusually prolonged duration, lasting for over 10 h from 17:00 UT on May 10 to 03:30 UT on May 11, 2024. This extended duration may be attributed to the continuous triggering of LSTIDs by auroral energy input during the geomagnetic storm. Additionally, significant differences in propagation characteristics, including velocities, azimuths, wavelengths, and traveling distances of LSTIDs, were observed between the two regions. These disparities in LSTID parameters are likely due to variations in the magnitude of energy input in the polar regions and local time differences in North America (14:00 LT) and Europe (19:00 LT), which cause diurnal electron-density contrast to influence LSTID propagation.
{"title":"Large-Scale Traveling Ionospheric Disturbances Over North America and Europe During the May 2024 Extreme Geomagnetic Storm","authors":"Long Tang;Hong Zhang;Yumei Li;Fan Xu;Fang Zou","doi":"10.1109/LGRS.2025.3608704","DOIUrl":"https://doi.org/10.1109/LGRS.2025.3608704","url":null,"abstract":"This study investigates the large-scale ionospheric traveling disturbances (LSTIDs) over North America and Europe associated with the intense geomagnetic storm in May 2024, utilizing total electron content (TEC) data derived from ground-based Global Navigation Satellite System (GNSS) stations. The findings reveal that the observed LSTIDs in both regions exhibited an unusually prolonged duration, lasting for over 10 h from 17:00 UT on May 10 to 03:30 UT on May 11, 2024. This extended duration may be attributed to the continuous triggering of LSTIDs by auroral energy input during the geomagnetic storm. Additionally, significant differences in propagation characteristics, including velocities, azimuths, wavelengths, and traveling distances of LSTIDs, were observed between the two regions. These disparities in LSTID parameters are likely due to variations in the magnitude of energy input in the polar regions and local time differences in North America (14:00 LT) and Europe (19:00 LT), which cause diurnal electron-density contrast to influence LSTID propagation.","PeriodicalId":91017,"journal":{"name":"IEEE geoscience and remote sensing letters : a publication of the IEEE Geoscience and Remote Sensing Society","volume":"22 ","pages":"1-5"},"PeriodicalIF":4.4,"publicationDate":"2025-09-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145090175","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2025-09-10DOI: 10.1109/LGRS.2025.3608489
Xiaofei Yu;Jie Ma;Liqiang Qiao
Remote sensing image change captioning (RSICC) is a challenging task that involves describing surface changes between bitemporal or multitemporal satellite images using natural language. This task requires both fine-grained visual understanding and expressive language generation. Transformer-based and long short-term memory (LSTM)-based models have shown promising results in this domain. However, they may encounter difficulties in generating flexible and diverse captions, particularly when training data are limited or imbalanced. While diffusion models provide richer textual outputs, they are often constrained by long inference times. To address these issues, we propose a novel diffusion-based framework, KD-RSCC, for efficient and expressive remote sensing change captioning. This framework utilizes the Karras sampling method to significantly reduce the number of steps required during inference, while preserving the quality and diversity of the generated captions. In addition, we introduce a large language model (LLM)-based evaluation strategy $text {G-Eval}_{text {RSCC}}$ to conduct a more comprehensive assessment of the semantic accuracy, fluency, and linguistic diversity of the generated descriptions. Experimental results demonstrate that KD-RSCC achieves an optimal balance between generation quality and inference speed, enhancing the flexibility and readability of its outputs. The code and supplementary materials are available at https://github.com/Fay-Y/KD_RSCC
{"title":"KD-RSCC: A Karras Diffusion Framework for Efficient Remote Sensing Change Captioning","authors":"Xiaofei Yu;Jie Ma;Liqiang Qiao","doi":"10.1109/LGRS.2025.3608489","DOIUrl":"https://doi.org/10.1109/LGRS.2025.3608489","url":null,"abstract":"Remote sensing image change captioning (RSICC) is a challenging task that involves describing surface changes between bitemporal or multitemporal satellite images using natural language. This task requires both fine-grained visual understanding and expressive language generation. Transformer-based and long short-term memory (LSTM)-based models have shown promising results in this domain. However, they may encounter difficulties in generating flexible and diverse captions, particularly when training data are limited or imbalanced. While diffusion models provide richer textual outputs, they are often constrained by long inference times. To address these issues, we propose a novel diffusion-based framework, KD-RSCC, for efficient and expressive remote sensing change captioning. This framework utilizes the Karras sampling method to significantly reduce the number of steps required during inference, while preserving the quality and diversity of the generated captions. In addition, we introduce a large language model (LLM)-based evaluation strategy <inline-formula> <tex-math>$text {G-Eval}_{text {RSCC}}$ </tex-math></inline-formula> to conduct a more comprehensive assessment of the semantic accuracy, fluency, and linguistic diversity of the generated descriptions. Experimental results demonstrate that KD-RSCC achieves an optimal balance between generation quality and inference speed, enhancing the flexibility and readability of its outputs. The code and supplementary materials are available at <uri>https://github.com/Fay-Y/KD_RSCC</uri>","PeriodicalId":91017,"journal":{"name":"IEEE geoscience and remote sensing letters : a publication of the IEEE Geoscience and Remote Sensing Society","volume":"22 ","pages":"1-5"},"PeriodicalIF":4.4,"publicationDate":"2025-09-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145090174","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2025-09-09DOI: 10.1109/LGRS.2025.3607840
Xiaosheng Yu;Weiqi Bai;Jubo Chen;Jiawei Huang;Zhuoqun Fang;Zhaokui Li
Accurate segmentation of very high-resolution remote sensing images is vital for downstream tasks. Most semantic segmentation methods fail to fully consider the inherent characteristics of the images, such as intricate backgrounds, significant intraclass variance, and spatial interdependence of geographic object distribution. To address these challenges, we propose an efficient global–local scene awareness network with rotary position embedding (RoGLSNet). Specifically, we introduce the dynamic global filter (DGF) module to adaptively select frequency components, thereby mitigating interference from background noise. For high intraclass variance, the class center aware block (CCAB) performs class-level contextual modeling with spatial information integration. Additionally, the rotary position embedding (RoPE) is incorporated into vanilla attention to indirectly model the positional and distance relationships of geographic target objects. Extensive experimental results on two widely used datasets demonstrate that RoGLSNet outperforms the state-of-the-art (SOTA) segmentation methods. The code is available at https://github.com/bai101315/RoGLSNet
{"title":"RoGLSNet: An Efficient Global–Local Scene Awareness Network With Rotary Position Embedding for Remote Image Segmentation","authors":"Xiaosheng Yu;Weiqi Bai;Jubo Chen;Jiawei Huang;Zhuoqun Fang;Zhaokui Li","doi":"10.1109/LGRS.2025.3607840","DOIUrl":"https://doi.org/10.1109/LGRS.2025.3607840","url":null,"abstract":"Accurate segmentation of very high-resolution remote sensing images is vital for downstream tasks. Most semantic segmentation methods fail to fully consider the inherent characteristics of the images, such as intricate backgrounds, significant intraclass variance, and spatial interdependence of geographic object distribution. To address these challenges, we propose an efficient global–local scene awareness network with rotary position embedding (RoGLSNet). Specifically, we introduce the dynamic global filter (DGF) module to adaptively select frequency components, thereby mitigating interference from background noise. For high intraclass variance, the class center aware block (CCAB) performs class-level contextual modeling with spatial information integration. Additionally, the rotary position embedding (RoPE) is incorporated into vanilla attention to indirectly model the positional and distance relationships of geographic target objects. Extensive experimental results on two widely used datasets demonstrate that RoGLSNet outperforms the state-of-the-art (SOTA) segmentation methods. The code is available at <uri>https://github.com/bai101315/RoGLSNet</uri>","PeriodicalId":91017,"journal":{"name":"IEEE geoscience and remote sensing letters : a publication of the IEEE Geoscience and Remote Sensing Society","volume":"22 ","pages":"1-5"},"PeriodicalIF":4.4,"publicationDate":"2025-09-09","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"145073326","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}