Hybrid-Module Transformer: enhancing speech emotion recognition with HuBERT, LSTM, and ResNet-50
Speech emotion recognition (SER) is a challenging task that involves identifying human emotions from speech. Traditional sequence models like recurrent neural network (RNN) and long short-term memory (LSTM) are limited by vanishing gradients and difficulty in capturing long-range dependencies. This...
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Artigo |
| Lingua: | Inglês |
| Pubblicazione: |
PeerJ Inc.
2025-10-01
|
| Serie: | PeerJ Computer Science |
| Soggetti: | |
| Accesso online: | https://peerj.com/articles/cs-3292.pdf |
| Tags: |
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
