QR kód

TASTA: Text‐Assisted Spatial and Temporal Attention Network for Video Question Answering

Video question answering (VideoQA) is a typical task that integrates language and vision. The key for VideoQA is to extract relevant and effective visual information for answering a specific question. Information selection is believed to be necessary for this task due to the large amount of irreleva...

Celý popis

Uloženo v:
Podrobná bibliografie
Hlavní autoři: Tian Wang, Boyao Hou, Jiakun Li, Peng Shi, Baochang Zhang, Hichem Snoussi
Médium: Artigo
Jazyk:Inglês
Vydáno: Wiley 2023-04-01
Edice:Advanced Intelligent Systems
Témata:
On-line přístup:https://doi.org/10.1002/aisy.202200131
Tagy: Přidat tag
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!