GLFormer: a hierarchical cross-layer transformer decoding framework with global–local feature collaboration for video action recognition
Abstract CLIP’s effectiveness across image recognition tasks stems from its large-scale multimodal pretraining, which equips it with powerful and transferable visual representations. However, when directly applied to video action recognition, the CLIP image encoder often falls short in modeling glob...
Сохранить в:
| Главные авторы: | , , , |
|---|---|
| Формат: | Artigo |
| Язык: | Inglês |
| Опубликовано: |
Springer
2026-04-01
|
| Серии: | Complex & Intelligent Systems |
| Предметы: | |
| Online-ссылка: | https://doi.org/10.1007/s40747-026-02314-3 |
| Метки: |
Нет меток, Требуется 1-ая метка записи!
|
