Multi-Module Co-Attention Model for Visual Question Answering
Visual Question Answering(VQA) is a typical multi-modal problem in computer vision and natural language processing.Most of the existing VQA models ignore the dynamic relationships of semantic information between two modes and the rich spatial structure of an image.For this reason, the paper proposes...
Na minha lista:
| Autor principal: | |
|---|---|
| Formato: | Artigo |
| Idioma: | Inglês |
| Publicado em: |
Editorial Office of Computer Engineering
2022-02-01
|
| Colecção: | Jisuanji gongcheng |
| Assuntos: | |
| Acesso em linha: | https://www.ecice06.com/fileup/1000-3428/PDF/20220233.pdf |
| Tags: |
Sem tags, seja o primeiro a adicionar uma tag!
|
