Côd QR

Structured Scene Parsing with a Hierarchical CLIP Model for Images

Visual Relationship Prediction (VRP) is crucial for advancing structured scene understanding, yet existing methods struggle with ineffective multimodal fusion, static relationship representations, and a lack of logical consistency. To address these limitations, this paper proposes a Hierarchical CLI...

Disgrifiad llawn

Wedi'i Gadw mewn:
Manylion Llyfryddiaeth
Prif Awduron: Yunhao Sun, Xiaoao Chen, Heng Chen, Yiduo Liang, Ruihua Qi
Fformat: Artigo
Iaith:Inglês
Cyhoeddwyd: MDPI AG 2026-01-01
Cyfres:Applied Sciences
Pynciau:
Mynediad Ar-lein:https://www.mdpi.com/2076-3417/16/2/788
Tagiau: Ychwanegu Tag
Dim Tagiau, Byddwch y cyntaf i dagio'r cofnod hwn!