Interactive text-guided image segmentation via vision Mamba and large language models
Abstract This paper proposes an H-type Bidirectional Alignment Network for text-guided image segmentation, enabling efficient and accurate cross-modal feature fusion. For visual encoding, a 12-layer Vision Mamba with four stages is adopted. It captures long-range dependencies via a Selective State S...
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Artigo |
| Lingua: | Inglês |
| Pubblicazione: |
Nature Portfolio
2026-03-01
|
| Serie: | Scientific Reports |
| Soggetti: | |
| Accesso online: | https://doi.org/10.1038/s41598-026-43841-w |
| Tags: |
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
