Context-aware Image Understanding in VQA with Dense-captioning
Visual Question Answering (VQA) is a complex task that requires models to jointly analyze visual and textual inputs to generate accurate answers. Reasoning and inference are critical for addressing questions that involve relationships, spatial arrangements, and contextual details within an image. In...
Na minha lista:
| Principais autores: | , , |
|---|---|
| Formato: | Artigo |
| Idioma: | Inglês |
| Publicado em: |
Iran Telecom Research Center
2025-09-01
|
| coleção: | International Journal of Information and Communication Technology Research |
| Assuntos: | |
| Acesso em linha: | http://ijict.itrc.ac.ir/article-1-740-en.pdf |
| Tags: |
Sem tags, seja o primeiro a adicionar uma tag!
|
