QR Code (код быстрого отклика)

Multi-Module Co-Attention Model for Visual Question Answering

Visual Question Answering(VQA) is a typical multi-modal problem in computer vision and natural language processing.Most of the existing VQA models ignore the dynamic relationships of semantic information between two modes and the rich spatial structure of an image.For this reason, the paper proposes...

Полное описание

Сохранить в:
Библиографические подробности
Главный автор: ZOU Pinrong, XIAO Feng, ZHANG Wenjuan, ZHANG Wanyu, WANG Chenyang
Формат: Artigo
Язык:Inglês
Опубликовано: Editorial Office of Computer Engineering 2022-02-01
Серии:Jisuanji gongcheng
Предметы:
Online-ссылка:https://www.ecice06.com/fileup/1000-3428/PDF/20220233.pdf
Метки: Добавить метку
Нет меток, Требуется 1-ая метка записи!