EmoBridge: Aligning Speech and Language for Emotion Recognition via Q-Former
Speech emotion recognition (SER) remains a challenging task due to the limited affective cues in unimodal representations and the difficulty of aligning heterogeneous features in multimodal systems. Although multimodal large language models (MLLMs) have recently achieved notable progress in affectiv...
שמור ב:
| Principais autores: | , , |
|---|---|
| פורמט: | Artigo |
| שפה: | Inglês |
| יצא לאור: |
IEEE
2025-01-01
|
| סדרה: | IEEE Access |
| נושאים: | |
| גישה מקוונת: | https://ieeexplore.ieee.org/document/11263794/ |
| תגים: |
אין תגיות, היה/י הראשונ/ה לתייג את הרשומה!
|
