QR kód

BERTIVITS: The Posterior Encoder Fusion of Pre-Trained Models and Residual Skip Connections for End-to-End Speech Synthesis

Enhancing the naturalness and rhythmicity of generated audio in end-to-end speech synthesis is crucial. The current state-of-the-art (SOTA) model, VITS, utilizes a conditional variational autoencoder architecture. However, it faces challenges, such as limited robustness, due to training solely on te...

Celý popis

Uloženo v:
Podrobná bibliografie
Hlavní autoři: Zirui Wang, Minqi Song, Dongbo Zhou
Médium: Artigo
Jazyk:Inglês
Vydáno: MDPI AG 2024-06-01
Edice:Applied Sciences
Témata:
On-line přístup:https://www.mdpi.com/2076-3417/14/12/5060
Tagy: Přidat tag
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!