QR kód

Image Description Generation Method by Panoptic Segmentation and Multi-Visual-Feature Fusion

Due to their powerful sequence modeling capabilities, Transformer-based image captioning models have demonstrated remarkable performance. However, most of these models typically utilize region visual features to perform encoding and decoding, which cannot fully use the fine-grained information of th...

Celý popis

Uloženo v:
Podrobná bibliografie
Hlavní autor: LIU Mingming, LU Jinfu, LIU Hao, ZHANG Haiyan
Médium: Artigo
Jazyk:Inglês
Vydáno: Editorial Office of Computer Engineering 2024-11-01
Edice:Jisuanji gongcheng
Témata:
On-line přístup:https://www.ecice06.com/fileup/1000-3428/PDF/20241129.pdf
Tagy: Přidat tag
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!