QR kód

Lipreading Architecture Based on Multiple Convolutional Neural Networks for Sentence-Level Visual Speech Recognition

In visual speech recognition (VSR), speech is transcribed using only visual information to interpret tongue and teeth movements. Recently, deep learning has shown outstanding performance in VSR, with accuracy exceeding that of lipreaders on benchmark datasets. However, several problems still exist w...

Celý popis

Uloženo v:
Podrobná bibliografie
Hlavní autoři: Sanghun Jeon, Ahmed Elsharkawy, Mun Sang Kim
Médium: Artigo
Jazyk:Inglês
Vydáno: MDPI AG 2021-12-01
Edice:Sensors
Témata:
On-line přístup:https://www.mdpi.com/1424-8220/22/1/72
Tagy: Přidat tag
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!