Structured Scene Parsing with a Hierarchical CLIP Model for Images
Visual Relationship Prediction (VRP) is crucial for advancing structured scene understanding, yet existing methods struggle with ineffective multimodal fusion, static relationship representations, and a lack of logical consistency. To address these limitations, this paper proposes a Hierarchical CLI...
Wedi'i Gadw mewn:
| Prif Awduron: | , , , , |
|---|---|
| Fformat: | Artigo |
| Iaith: | Inglês |
| Cyhoeddwyd: |
MDPI AG
2026-01-01
|
| Cyfres: | Applied Sciences |
| Pynciau: | |
| Mynediad Ar-lein: | https://www.mdpi.com/2076-3417/16/2/788 |
| Tagiau: |
Dim Tagiau, Byddwch y cyntaf i dagio'r cofnod hwn!
|
