Hybrid-Module Transformer: enhancing speech emotion recognition with HuBERT, LSTM, and ResNet-50
Speech emotion recognition (SER) is a challenging task that involves identifying human emotions from speech. Traditional sequence models like recurrent neural network (RNN) and long short-term memory (LSTM) are limited by vanishing gradients and difficulty in capturing long-range dependencies. This...
I tiakina i:
| Ngā kaituhi matua: | , , , |
|---|---|
| Hōputu: | Artigo |
| Reo: | Inglês |
| I whakaputaina: |
PeerJ Inc.
2025-10-01
|
| Rangatū: | PeerJ Computer Science |
| Ngā marau: | |
| Urunga tuihono: | https://peerj.com/articles/cs-3292.pdf |
| Ngā Tūtohu: |
Kāore He Tūtohu, Me noho koe te mea tuatahi ki te tūtohu i tēnei pūkete!
|
