GLFormer: a hierarchical cross-layer transformer decoding framework with global–local feature collaboration for video action recognition
Abstract CLIP’s effectiveness across image recognition tasks stems from its large-scale multimodal pretraining, which equips it with powerful and transferable visual representations. However, when directly applied to video action recognition, the CLIP image encoder often falls short in modeling glob...
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Artigo |
| Language: | Inglês |
| Published: |
Springer
2026-04-01
|
| Series: | Complex & Intelligent Systems |
| Subjects: | |
| Online Access: | https://doi.org/10.1007/s40747-026-02314-3 |
| Tags: |
No Tags, Be the first to tag this record!
|
