Co-LLaVA: Efficient Remote Sensing Visual Question Answering via Model Collaboration
Large vision language models (LVLMs) are built upon large language models (LLMs) and incorporate non-textual modalities; they can perform various multimodal tasks. Applying LVLMs in remote sensing (RS) visual question answering (VQA) tasks can take advantage of the powerful capabilities to promote t...
保存先:
| 主要な著者: | , , , , , |
|---|---|
| フォーマット: | Artigo |
| 言語: | Inglês |
| 出版事項: |
MDPI AG
2025-01-01
|
| シリーズ: | Remote Sensing |
| 主題: | |
| オンライン・アクセス: | https://www.mdpi.com/2072-4292/17/3/466 |
| タグ: |
タグなし, このレコードへの初めてのタグを付けませんか!
|
