CLIP-Guided Dynamic Feature Fusion With Relation-Enhanced Attention for Remote Sensing Image Captioning
Remote sensing image captioning is a multimodal task that aims to automatically generate descriptions for remote sensing images. However, the inherent scale diversity and semantic complexity of RSI pose significant challenges in semantic alignment and explicit modeling of complex scene structures. T...
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Artigo |
| Sprache: | Inglês |
| Veröffentlicht: |
IEEE
2026-01-01
|
| Schriftenreihe: | IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing |
| Schlagworte: | |
| Online-Zugang: | https://ieeexplore.ieee.org/document/11488868/ |
| Tags: |
Keine Tags, Fügen Sie das erste Tag hinzu!
|
