Código QR (código de barras bidimensional)

Structured Scene Parsing with a Hierarchical CLIP Model for Images

Visual Relationship Prediction (VRP) is crucial for advancing structured scene understanding, yet existing methods struggle with ineffective multimodal fusion, static relationship representations, and a lack of logical consistency. To address these limitations, this paper proposes a Hierarchical CLI...

תיאור מלא

שמור ב:
מידע ביבליוגרפי
Principais autores: Yunhao Sun, Xiaoao Chen, Heng Chen, Yiduo Liang, Ruihua Qi
פורמט: Artigo
שפה:Inglês
יצא לאור: MDPI AG 2026-01-01
סדרה:Applied Sciences
נושאים:
גישה מקוונת:https://www.mdpi.com/2076-3417/16/2/788
תגים: הוספת תג
אין תגיות, היה/י הראשונ/ה לתייג את הרשומה!