End-to-End Multi-Modal Speaker Change Detection with Pre-Trained Models
In this work, we propose a multi-modal speaker change detection (SCD) approach with focal loss, which integrates both audio and text features to enhance detection performance. The proposed approach utilizes pre-trained large-scale models for feature extraction and incorporates a self-attention mecha...
Αποθηκεύτηκε σε:
| Κύριοι συγγραφείς: | , , , , |
|---|---|
| Μορφή: | Artigo |
| Γλώσσα: | Inglês |
| Έκδοση: |
MDPI AG
2025-04-01
|
| Σειρά: | Applied Sciences |
| Θέματα: | |
| Διαθέσιμο Online: | https://www.mdpi.com/2076-3417/15/8/4324 |
| Ετικέτες: |
Δεν υπάρχουν, Καταχωρήστε ετικέτα πρώτοι!
|
