QR код

Hybrid-Module Transformer: enhancing speech emotion recognition with HuBERT, LSTM, and ResNet-50

Speech emotion recognition (SER) is a challenging task that involves identifying human emotions from speech. Traditional sequence models like recurrent neural network (RNN) and long short-term memory (LSTM) are limited by vanishing gradients and difficulty in capturing long-range dependencies. This...

Повний опис

Збережено в:
Бібліографічні деталі
Автори: Xindong Huang, Wuhui Lin, Maming Chen, Hua Shi
Формат: Artigo
Мова:Inglês
Опубліковано: PeerJ Inc. 2025-10-01
Серія:PeerJ Computer Science
Предмети:
Онлайн доступ:https://peerj.com/articles/cs-3292.pdf
Теги: Додати тег
Немає тегів, Будьте першим, хто поставить тег для цього запису!