Código QR (código de barras bidimensional)

LLaVA-Pose: Keypoint-Integrated Instruction Tuning for Human Pose and Action Understanding

Current vision–language models (VLMs) are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized vision–language instruction-following data. We introduce a method for...

Fuld beskrivelse

Na minha lista:
Bibliografiske detaljer
Principais autores: Dewen Zhang, Tahir Hussain, Wangpeng An, Hayaru Shouno
Format: Artigo
Sprog:Inglês
Udgivet: MDPI AG 2025-08-01
Serier:Sensors
Fag:
Online adgang:https://www.mdpi.com/1424-8220/25/16/5213
Tags: Tilføj Tag
Ingen Tags, Vær først til at tagge denne postø!