TAME: Temporal-Aware Mixture-of-Experts for Text–Video Retrieval
Text–Video Retrieval (TVR) retrieves videos that match a natural-language query, but extending image–text models such as CLIP to videos is fundamentally limited by the lack of temporal modeling. Videos exhibit frame-wise heterogeneity in appearance and motion, and compressing all frame...
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Artigo |
| Lingua: | Inglês |
| Pubblicazione: |
IEEE
2026-01-01
|
| Serie: | IEEE Access |
| Soggetti: | |
| Accesso online: | https://ieeexplore.ieee.org/document/11364210/ |
| Tags: |
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
