Interactive text-guided image segmentation via vision Mamba and large language models
Abstract This paper proposes an H-type Bidirectional Alignment Network for text-guided image segmentation, enabling efficient and accurate cross-modal feature fusion. For visual encoding, a 12-layer Vision Mamba with four stages is adopted. It captures long-range dependencies via a Selective State S...
Uloženo v:
| Hlavní autoři: | , , |
|---|---|
| Médium: | Artigo |
| Jazyk: | Inglês |
| Vydáno: |
Nature Portfolio
2026-03-01
|
| Edice: | Scientific Reports |
| Témata: | |
| On-line přístup: | https://doi.org/10.1038/s41598-026-43841-w |
| Tagy: |
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!
|
