Nanami: Hybrid Embedding With Voice Large Language Models for Audio Retrieval
Cross-modal retrieval has become essential in establishing semantic correspondences between heterogeneous data modalities, particularly in text-audio retrieval applications. Generally, current contrastive learning-based approaches (e.g., CLAP) perform well in audio retrieval with short descriptive c...
Guardat en:
| Autor principal: | |
|---|---|
| Format: | Artigo |
| Idioma: | Inglês |
| Publicat: |
IEEE
2025-01-01
|
| Col·lecció: | IEEE Access |
| Matèries: | |
| Accés en línia: | https://ieeexplore.ieee.org/document/11224486/ |
| Etiquetes: |
Sense etiquetes, Sigues el primer a etiquetar aquest registre!
|
