QR Code (код быстрого отклика)

GLFormer: a hierarchical cross-layer transformer decoding framework with global–local feature collaboration for video action recognition

Abstract CLIP’s effectiveness across image recognition tasks stems from its large-scale multimodal pretraining, which equips it with powerful and transferable visual representations. However, when directly applied to video action recognition, the CLIP image encoder often falls short in modeling glob...

Полное описание

Сохранить в:
Библиографические подробности
Главные авторы: Hanbo Wu, Xin Ma, Xiang Li, Rui Song
Формат: Artigo
Язык:Inglês
Опубликовано: Springer 2026-04-01
Серии:Complex & Intelligent Systems
Предметы:
Online-ссылка:https://doi.org/10.1007/s40747-026-02314-3
Метки: Добавить метку
Нет меток, Требуется 1-ая метка записи!