TASTA: Text‐Assisted Spatial and Temporal Attention Network for Video Question Answering
Video question answering (VideoQA) is a typical task that integrates language and vision. The key for VideoQA is to extract relevant and effective visual information for answering a specific question. Information selection is believed to be necessary for this task due to the large amount of irreleva...
Uloženo v:
| Hlavní autoři: | , , , , , |
|---|---|
| Médium: | Artigo |
| Jazyk: | Inglês |
| Vydáno: |
Wiley
2023-04-01
|
| Edice: | Advanced Intelligent Systems |
| Témata: | |
| On-line přístup: | https://doi.org/10.1002/aisy.202200131 |
| Tagy: |
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!
|
