Optimizing feature fusion for improved zero-shot adaptation in text-to-speech synthesis
Abstract In the era of advanced text-to-speech (TTS) systems capable of generating high-fidelity, human-like speech by referring a reference speech, voice cloning (VC), or zero-shot TTS (ZS-TTS), stands out as an important subtask. A primary challenge in VC is maintaining speech quality and speaker...
I tiakina i:
| Ngā kaituhi matua: | , , , , |
|---|---|
| Hōputu: | Artigo |
| Reo: | Inglês |
| I whakaputaina: |
SpringerOpen
2024-05-01
|
| Rangatū: | EURASIP Journal on Audio, Speech, and Music Processing |
| Ngā marau: | |
| Urunga tuihono: | https://doi.org/10.1186/s13636-024-00351-9 |
| Ngā Tūtohu: |
Kāore He Tūtohu, Me noho koe te mea tuatahi ki te tūtohu i tēnei pūkete!
|
