Código QR

V-PRUNE: Semantic-Aware Patch Pruning Before Tokenization in Vision–Language Model Inference

Recent vision–language models (VLMs) achieve strong performance across multimodal benchmarks but suffer from high inference costs due to the large number of visual tokens. Prior studies have shown that many image tokens receive consistently low attention scores during inference, indicating that a su...

ver descrição completa

Na minha lista:
Detalhes bibliográficos
Principais autores: Hyein Seo, Yong Suk Choi
Formato: Artigo
Idioma:Inglês
Publicado em: MDPI AG 2025-08-01
Colecção:Applied Sciences
Assuntos:
Acesso em linha:https://www.mdpi.com/2076-3417/15/17/9463
Tags: Adicionar Tag
Sem tags, seja o primeiro a adicionar uma tag!