Hybrid-Module Transformer: enhancing speech emotion recognition with HuBERT, LSTM, and ResNet-50
Speech emotion recognition (SER) is a challenging task that involves identifying human emotions from speech. Traditional sequence models like recurrent neural network (RNN) and long short-term memory (LSTM) are limited by vanishing gradients and difficulty in capturing long-range dependencies. This...
Збережено в:
| Автори: | , , , |
|---|---|
| Формат: | Artigo |
| Мова: | Inglês |
| Опубліковано: |
PeerJ Inc.
2025-10-01
|
| Серія: | PeerJ Computer Science |
| Предмети: | |
| Онлайн доступ: | https://peerj.com/articles/cs-3292.pdf |
| Теги: |
Немає тегів, Будьте першим, хто поставить тег для цього запису!
|
