WMSA–WBS: Efficient Wave Multi-Head Self-Attention with Wavelet Bottleneck
The critical component of the vision transformer (ViT) architecture is multi-head self-attention (MSA), which enables the encoding of long-range dependencies and heterogeneous interactions. However, MSA has two significant limitations: its limited ability to capture local features and its high compu...
Bewaard in:
| Hoofdauteurs: | , , , |
|---|---|
| Formaat: | Artigo |
| Taal: | Inglês |
| Gepubliceerd in: |
MDPI AG
2025-08-01
|
| Reeks: | Sensors |
| Onderwerpen: | |
| Online toegang: | https://www.mdpi.com/1424-8220/25/16/5046 |
| Tags: |
Geen labels, Wees de eerste die dit record labelt!
|
