QR код

Research on Text Representation of Video Content Based on Multi-Modal Fusion and Multi-Layer Attention

Aiming at the challenges of single-text representation and low accuracy of existing video content text-representation models, a video content text-reprsentation model that integrates frame-level image and audio information is proposed.The network structure of the model includes a single-mode embeddi...

Повний опис

Збережено в:
Бібліографічні деталі
Автор: ZHAO Hong, GUO Lan, CHEN Zhiwen, ZHENG Houze
Формат: Artigo
Мова:Inglês
Опубліковано: Editorial Office of Computer Engineering 2022-10-01
Серія:Jisuanji gongcheng
Предмети:
Онлайн доступ:https://www.ecice06.com/fileup/1000-3428/PDF/20221006.pdf
Теги: Додати тег
Немає тегів, Будьте першим, хто поставить тег для цього запису!