Nanami: Hybrid Embedding With Voice Large Language Models for Audio Retrieval
Cross-modal retrieval has become essential in establishing semantic correspondences between heterogeneous data modalities, particularly in text-audio retrieval applications. Generally, current contrastive learning-based approaches (e.g., CLAP) perform well in audio retrieval with short descriptive c...
Na minha lista:
| Autor principal: | |
|---|---|
| Formato: | Artigo |
| Idioma: | Inglês |
| Publicado em: |
IEEE
2025-01-01
|
| Colecção: | IEEE Access |
| Assuntos: | |
| Acesso em linha: | https://ieeexplore.ieee.org/document/11224486/ |
| Tags: |
Sem tags, seja o primeiro a adicionar uma tag!
|
