QRコード

Co-LLaVA: Efficient Remote Sensing Visual Question Answering via Model Collaboration

Large vision language models (LVLMs) are built upon large language models (LLMs) and incorporate non-textual modalities; they can perform various multimodal tasks. Applying LVLMs in remote sensing (RS) visual question answering (VQA) tasks can take advantage of the powerful capabilities to promote t...

詳細記述

保存先:
書誌詳細
主要な著者: Fan Liu, Wenwen Dai, Chuanyi Zhang, Jiale Zhu, Liang Yao, Xin Li
フォーマット: Artigo
言語:Inglês
出版事項: MDPI AG 2025-01-01
シリーズ:Remote Sensing
主題:
オンライン・アクセス:https://www.mdpi.com/2072-4292/17/3/466
タグ: タグ追加
タグなし, このレコードへの初めてのタグを付けませんか!