QR-koodi

Rethinking Visual Attention for Reducing Hallucination in Large Vision–Language Models

Large Vision–Language Models (LVLMs) have achieved strong performance in multimodal understanding and generation. However, they remain prone to hallucination, where generated content deviates from the visual input, reducing output reliability. We analyze the attention mechanism and identify two key...

Täydet tiedot

Tallennettuna:
Bibliografiset tiedot
Päätekijät: Xuewen Li, Yuan Liu
Aineistotyyppi: Artigo
Kieli:Inglês
Julkaistu: MDPI AG 2026-04-01
Sarja:Applied Sciences
Aiheet:
Linkit:https://www.mdpi.com/2076-3417/16/9/4143
Tagit: Lisää tagi
Ei tageja, Lisää ensimmäinen tagi!