LLaVA-Pose: Keypoint-Integrated Instruction Tuning for Human Pose and Action Understanding
Current vision–language models (VLMs) are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized vision–language instruction-following data. We introduce a method for...
Na minha lista:
| Principais autores: | , , , |
|---|---|
| Format: | Artigo |
| Sprog: | Inglês |
| Udgivet: |
MDPI AG
2025-08-01
|
| Serier: | Sensors |
| Fag: | |
| Online adgang: | https://www.mdpi.com/1424-8220/25/16/5213 |
| Tags: |
Ingen Tags, Vær først til at tagge denne postø!
|
