Código QR (código de barras bidimensional)

LLaVA-Pose: Keypoint-Integrated Instruction Tuning for Human Pose and Action Understanding

Current vision–language models (VLMs) are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized vision–language instruction-following data. We introduce a method for...

ver descrição completa

Na minha lista:
Detalhes bibliográficos
Principais autores: Dewen Zhang, Tahir Hussain, Wangpeng An, Hayaru Shouno
Formato: Artigo
Idioma:Inglês
Publicado em: MDPI AG 2025-08-01
coleção:Sensors
Assuntos:
Acesso em linha:https://www.mdpi.com/1424-8220/25/16/5213
Tags: Adicionar Tag
Sem tags, seja o primeiro a adicionar uma tag!