WMSA–WBS: Efficient Wave Multi-Head Self-Attention with Wavelet Bottleneck
The critical component of the vision transformer (ViT) architecture is multi-head self-attention (MSA), which enables the encoding of long-range dependencies and heterogeneous interactions. However, MSA has two significant limitations: its limited ability to capture local features and its high compu...
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Artigo |
| Lenguaje: | Inglês |
| Publicado: |
MDPI AG
2025-08-01
|
| Colección: | Sensors |
| Materias: | |
| Acceso en línea: | https://www.mdpi.com/1424-8220/25/16/5046 |
| Etiquetas: |
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
