Research on Text Representation of Video Content Based on Multi-Modal Fusion and Multi-Layer Attention
Aiming at the challenges of single-text representation and low accuracy of existing video content text-representation models, a video content text-reprsentation model that integrates frame-level image and audio information is proposed.The network structure of the model includes a single-mode embeddi...
Збережено в:
| Автор: | |
|---|---|
| Формат: | Artigo |
| Мова: | Inglês |
| Опубліковано: |
Editorial Office of Computer Engineering
2022-10-01
|
| Серія: | Jisuanji gongcheng |
| Предмети: | |
| Онлайн доступ: | https://www.ecice06.com/fileup/1000-3428/PDF/20221006.pdf |
| Теги: |
Немає тегів, Будьте першим, хто поставить тег для цього запису!
|
