QR kód

Audio-Language Datasets of Scenes and Events: A Survey

Audio-language models (ALMs) generate linguistic descriptions of sound-producing events and scenes. Advances in dataset creation and computational power have led to significant progress in this domain. This paper surveys 69 datasets used to train ALMs, covering research up to September 2024 (<uri>ht...

Celý popis

Uloženo v:
Podrobná bibliografie
Hlavní autoři: Gijs Wijngaard, Elia Formisano, Michele Esposito, Michel Dumontier
Médium: Artigo
Jazyk:Inglês
Vydáno: IEEE 2025-01-01
Edice:IEEE Access
Témata:
On-line přístup:https://ieeexplore.ieee.org/document/10854210/
Tagy: Přidat tag
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!