Multimodal Fine-Tuning of LLMs for Robust Document Visual Question Answering
Document Visual Question Answering (DocVQA) necessitates comprehension of both the spatial layout and the textual content. Multimodal pretraining is a foundational component of existing vision-language models, including LayoutLM. However, they frequently lack integration with potent Large Language M...
Sparad:
| Huvudupphov: | , , , |
|---|---|
| Materialtyp: | Artigo |
| Språk: | Inglês |
| Utgiven: |
IEEE
2025-01-01
|
| Serie: | IEEE Access |
| Ämnen: | |
| Länkar: | https://ieeexplore.ieee.org/document/11184117/ |
| Taggar: |
Inga taggar, Lägg till första taggen!
|
