Cap4Bridge: Caption-Guided Cross-Modal Contextualization With Stochastic Augmentation for Text-Video Retrieval
A key challenge in text-video retrieval is bridging the semantic gap between information-rich videos and concise text queries. Existing methods often address this by incorporating auxiliary captions from Large Language Models (LLMs) or employing stochastic modeling. However, these approaches face cr...
Պահպանված է:
| Հիմնական հեղինակներ: | , , , , , |
|---|---|
| Ձևաչափ: | Artigo |
| Լեզու: | Inglês |
| Հրապարակվել է: |
IEEE
2026-01-01
|
| Շարք: | IEEE Access |
| Խորագրեր: | |
| Առցանց հասանելիություն: | https://ieeexplore.ieee.org/document/11474843/ |
| Ցուցիչներ: |
Չկան պիտակներ, Եղեք առաջինը, ով նշում է այս գրառումը!
|
