Cap4Bridge: Caption-Guided Cross-Modal Contextualization With Stochastic Augmentation for Text-Video Retrieval
A key challenge in text-video retrieval is bridging the semantic gap between information-rich videos and concise text queries. Existing methods often address this by incorporating auxiliary captions from Large Language Models (LLMs) or employing stochastic modeling. However, these approaches face cr...
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Artigo |
| Langue: | Inglês |
| Publié: |
IEEE
2026-01-01
|
| Collection: | IEEE Access |
| Sujets: | |
| Accès en ligne: | https://ieeexplore.ieee.org/document/11474843/ |
| Tags: |
Pas de tags, Soyez le premier à ajouter un tag!
|
