HATNet: Hierarchical Attention Transformer With RS-CLIP Patch Tokens for Remote Sensing Image Captioning
Remote sensing image captioning (RSIC) aims to generate natural language descriptions of critical visual content in overhead-view remote sensing images. However, existing methods often produce descriptions with significant omissions and misidentifications of scene types or object counts, particularl...
محفوظ في:
| المؤلفون الرئيسيون: | , , , , , , , |
|---|---|
| التنسيق: | Artigo |
| اللغة: | Inglês |
| منشور في: |
IEEE
2025-01-01
|
| سلاسل: | IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing |
| الموضوعات: | |
| الوصول للمادة أونلاين: | https://ieeexplore.ieee.org/document/11214231/ |
| الوسوم: |
لا توجد وسوم, كن أول من يضع وسما على هذه التسجيلة!
|
