Nanami: Hybrid Embedding With Voice Large Language Models for Audio Retrieval
Cross-modal retrieval has become essential in establishing semantic correspondences between heterogeneous data modalities, particularly in text-audio retrieval applications. Generally, current contrastive learning-based approaches (e.g., CLAP) perform well in audio retrieval with short descriptive c...
Tallennettuna:
| Päätekijä: | |
|---|---|
| Aineistotyyppi: | Artigo |
| Kieli: | Inglês |
| Julkaistu: |
IEEE
2025-01-01
|
| Sarja: | IEEE Access |
| Aiheet: | |
| Linkit: | https://ieeexplore.ieee.org/document/11224486/ |
| Tagit: |
Ei tageja, Lisää ensimmäinen tagi!
|
