Hybrid-Module Transformer: enhancing speech emotion recognition with HuBERT, LSTM, and ResNet-50
Speech emotion recognition (SER) is a challenging task that involves identifying human emotions from speech. Traditional sequence models like recurrent neural network (RNN) and long short-term memory (LSTM) are limited by vanishing gradients and difficulty in capturing long-range dependencies. This...
Furkejuvvon:
| Váldodahkkit: | , , , |
|---|---|
| Materiálatiipa: | Artigo |
| Giella: | Inglês |
| Almmustuhtton: |
PeerJ Inc.
2025-10-01
|
| Ráidu: | PeerJ Computer Science |
| Fáttát: | |
| Liŋkkat: | https://peerj.com/articles/cs-3292.pdf |
| Fáddágilkorat: |
Eai fáddágilkorat, Lasit vuosttaš fáddágilkora!
|
