BERTIVITS: The Posterior Encoder Fusion of Pre-Trained Models and Residual Skip Connections for End-to-End Speech Synthesis
Enhancing the naturalness and rhythmicity of generated audio in end-to-end speech synthesis is crucial. The current state-of-the-art (SOTA) model, VITS, utilizes a conditional variational autoencoder architecture. However, it faces challenges, such as limited robustness, due to training solely on te...
Uloženo v:
| Hlavní autoři: | , , |
|---|---|
| Médium: | Artigo |
| Jazyk: | Inglês |
| Vydáno: |
MDPI AG
2024-06-01
|
| Edice: | Applied Sciences |
| Témata: | |
| On-line přístup: | https://www.mdpi.com/2076-3417/14/12/5060 |
| Tagy: |
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!
|
