QR Code

GLFormer: a hierarchical cross-layer transformer decoding framework with global–local feature collaboration for video action recognition

Abstract CLIP’s effectiveness across image recognition tasks stems from its large-scale multimodal pretraining, which equips it with powerful and transferable visual representations. However, when directly applied to video action recognition, the CLIP image encoder often falls short in modeling glob...

Full description

Saved in:
Bibliographic Details
Main Authors: Hanbo Wu, Xin Ma, Xiang Li, Rui Song
Format: Artigo
Language:Inglês
Published: Springer 2026-04-01
Series:Complex & Intelligent Systems
Subjects:
Online Access:https://doi.org/10.1007/s40747-026-02314-3
Tags: Add Tag
No Tags, Be the first to tag this record!