AFT-SAM: Adaptive Fusion Transformer with a Sparse Attention Mechanism for Audio–Visual Speech Recognition
Aiming at the problems of serious information redundancy, complex inter-modal information interaction, and difficult multimodal fusion faced by the audio–visual speech recognition system when dealing with complex multimodal information, this paper proposes an adaptive fusion transformer algorithm (A...
Sábháilte in:
| Príomhchruthaitheoirí: | , , , , |
|---|---|
| Formáid: | Artigo |
| Teanga: | Inglês |
| Foilsithe / Cruthaithe: |
MDPI AG
2024-12-01
|
| Sraith: | Applied Sciences |
| Ábhair: | |
| Rochtain ar líne: | https://www.mdpi.com/2076-3417/15/1/199 |
| Clibeanna: |
Níl clibeanna ann, Bí ar an gcéad duine le clib a chur leis an taifead seo!
|
