Nanami: Hybrid Embedding With Voice Large Language Models for Audio Retrieval
Cross-modal retrieval has become essential in establishing semantic correspondences between heterogeneous data modalities, particularly in text-audio retrieval applications. Generally, current contrastive learning-based approaches (e.g., CLAP) perform well in audio retrieval with short descriptive c...
Furkejuvvon:
| Váldodahkki: | |
|---|---|
| Materiálatiipa: | Artigo |
| Giella: | Inglês |
| Almmustuhtton: |
IEEE
2025-01-01
|
| Ráidu: | IEEE Access |
| Fáttát: | |
| Liŋkkat: | https://ieeexplore.ieee.org/document/11224486/ |
| Fáddágilkorat: |
Eai fáddágilkorat, Lasit vuosttaš fáddágilkora!
|
