Semantic-Aligned Cross-Modal Visual Grounding Network with Transformers
Multi-modal deep learning methods have achieved great improvements in visual grounding; their objective is to localize text-specified objects in images. Most of the existing methods can localize and classify objects with significant appearance differences but suffer from the misclassification proble...
Na minha lista:
| Principais autores: | , |
|---|---|
| Format: | Artigo |
| Sprog: | Inglês |
| Udgivet: |
MDPI AG
2023-05-01
|
| Serier: | Applied Sciences |
| Fag: | |
| Online adgang: | https://www.mdpi.com/2076-3417/13/9/5649 |
| Tags: |
Ingen Tags, Vær først til at tagge denne postø!
|
