QR-kod

Multimodal Fine-Tuning of LLMs for Robust Document Visual Question Answering

Document Visual Question Answering (DocVQA) necessitates comprehension of both the spatial layout and the textual content. Multimodal pretraining is a foundational component of existing vision-language models, including LayoutLM. However, they frequently lack integration with potent Large Language M...

Full beskrivning

Sparad:
Bibliografiska uppgifter
Huvudupphov: Sahil Tripathi, Md Tabrez Nafis, Imran Hussain, Abdul Khader Jilani Saudagar
Materialtyp: Artigo
Språk:Inglês
Utgiven: IEEE 2025-01-01
Serie:IEEE Access
Ämnen:
Länkar:https://ieeexplore.ieee.org/document/11184117/
Taggar: Lägg till en tagg
Inga taggar, Lägg till första taggen!