Multi-Module Co-Attention Model for Visual Question Answering
Visual Question Answering(VQA) is a typical multi-modal problem in computer vision and natural language processing.Most of the existing VQA models ignore the dynamic relationships of semantic information between two modes and the rich spatial structure of an image.For this reason, the paper proposes...
Сохранить в:
| Главный автор: | |
|---|---|
| Формат: | Artigo |
| Язык: | Inglês |
| Опубликовано: |
Editorial Office of Computer Engineering
2022-02-01
|
| Серии: | Jisuanji gongcheng |
| Предметы: | |
| Online-ссылка: | https://www.ecice06.com/fileup/1000-3428/PDF/20220233.pdf |
| Метки: |
Нет меток, Требуется 1-ая метка записи!
|
