QR-Code

Rethinking Visual Attention for Reducing Hallucination in Large Vision–Language Models

Large Vision–Language Models (LVLMs) have achieved strong performance in multimodal understanding and generation. However, they remain prone to hallucination, where generated content deviates from the visual input, reducing output reliability. We analyze the attention mechanism and identify two key...

Ausführliche Beschreibung

Gespeichert in:
Bibliografische Detailangaben
Hauptverfasser: Xuewen Li, Yuan Liu
Format: Artigo
Sprache:Inglês
Veröffentlicht: MDPI AG 2026-04-01
Schriftenreihe:Applied Sciences
Schlagworte:
Online-Zugang:https://www.mdpi.com/2076-3417/16/9/4143
Tags: Tag hinzufügen
Keine Tags, Fügen Sie das erste Tag hinzu!