Gravar e-mail: Cap4Bridge: Caption-Guided Cross-Modal Contextualization With Stochastic Augmentation for Text-Video Retrieval