CLIP-Guided Dynamic Feature Fusion With Relation-Enhanced Attention for Remote Sensing Image Captioning
Remote sensing image captioning is a multimodal task that aims to automatically generate descriptions for remote sensing images. However, the inherent scale diversity and semantic complexity of RSI pose significant challenges in semantic alignment and explicit modeling of complex scene structures. T...
保存先:
| 主要な著者: | , , , , |
|---|---|
| フォーマット: | Artigo |
| 言語: | Inglês |
| 出版事項: |
IEEE
2026-01-01
|
| シリーズ: | IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing |
| 主題: | |
| オンライン・アクセス: | https://ieeexplore.ieee.org/document/11488868/ |
| タグ: |
タグなし, このレコードへの初めてのタグを付けませんか!
|
