QR kód

Interactive text-guided image segmentation via vision Mamba and large language models

Abstract This paper proposes an H-type Bidirectional Alignment Network for text-guided image segmentation, enabling efficient and accurate cross-modal feature fusion. For visual encoding, a 12-layer Vision Mamba with four stages is adopted. It captures long-range dependencies via a Selective State S...

Celý popis

Uloženo v:
Podrobná bibliografie
Hlavní autoři: Yao Meng, Haochen Sun, Wei Jiang
Médium: Artigo
Jazyk:Inglês
Vydáno: Nature Portfolio 2026-03-01
Edice:Scientific Reports
Témata:
On-line přístup:https://doi.org/10.1038/s41598-026-43841-w
Tagy: Přidat tag
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!