Multi-Module Co-Attention Model for Visual Question Answering
Visual Question Answering(VQA) is a typical multi-modal problem in computer vision and natural language processing.Most of the existing VQA models ignore the dynamic relationships of semantic information between two modes and the rich spatial structure of an image.For this reason, the paper proposes...
Na minha lista:
| Hovedforfatter: | |
|---|---|
| Format: | Artigo |
| Sprog: | Inglês |
| Udgivet: |
Editorial Office of Computer Engineering
2022-02-01
|
| Serier: | Jisuanji gongcheng |
| Fag: | |
| Online adgang: | https://www.ecice06.com/fileup/1000-3428/PDF/20220233.pdf |
| Tags: |
Ingen Tags, Vær først til at tagge denne postø!
|
