QR-Code

CLIP-Guided Dynamic Feature Fusion With Relation-Enhanced Attention for Remote Sensing Image Captioning

Remote sensing image captioning is a multimodal task that aims to automatically generate descriptions for remote sensing images. However, the inherent scale diversity and semantic complexity of RSI pose significant challenges in semantic alignment and explicit modeling of complex scene structures. T...

Ausführliche Beschreibung

Gespeichert in:
Bibliografische Detailangaben
Hauptverfasser: Haifeng Sima, Wenqing Jiang, JianLong Wang, Changchun Li, Mingliang Xu
Format: Artigo
Sprache:Inglês
Veröffentlicht: IEEE 2026-01-01
Schriftenreihe:IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing
Schlagworte:
Online-Zugang:https://ieeexplore.ieee.org/document/11488868/
Tags: Tag hinzufügen
Keine Tags, Fügen Sie das erste Tag hinzu!