TAME: Temporal-Aware Mixture-of-Experts for Text–Video Retrieval
Text–Video Retrieval (TVR) retrieves videos that match a natural-language query, but extending image–text models such as CLIP to videos is fundamentally limited by the lack of temporal modeling. Videos exhibit frame-wise heterogeneity in appearance and motion, and compressing all frame...
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Artigo |
| Sprache: | Inglês |
| Veröffentlicht: |
IEEE
2026-01-01
|
| Schriftenreihe: | IEEE Access |
| Schlagworte: | |
| Online-Zugang: | https://ieeexplore.ieee.org/document/11364210/ |
| Tags: |
Keine Tags, Fügen Sie das erste Tag hinzu!
|
