LLaVA-Pose: Keypoint-Integrated Instruction Tuning for Human Pose and Action Understanding
Current vision–language models (VLMs) are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized vision–language instruction-following data. We introduce a method for...
Na minha lista:
| Principais autores: | , , , |
|---|---|
| Formato: | Artigo |
| Idioma: | Inglês |
| Publicado em: |
MDPI AG
2025-08-01
|
| coleção: | Sensors |
| Assuntos: | |
| Acesso em linha: | https://www.mdpi.com/1424-8220/25/16/5213 |
| Tags: |
Sem tags, seja o primeiro a adicionar uma tag!
|
