Image Description Generation Method by Panoptic Segmentation and Multi-Visual-Feature Fusion
Due to their powerful sequence modeling capabilities, Transformer-based image captioning models have demonstrated remarkable performance. However, most of these models typically utilize region visual features to perform encoding and decoding, which cannot fully use the fine-grained information of th...
Uloženo v:
| Hlavní autor: | |
|---|---|
| Médium: | Artigo |
| Jazyk: | Inglês |
| Vydáno: |
Editorial Office of Computer Engineering
2024-11-01
|
| Edice: | Jisuanji gongcheng |
| Témata: | |
| On-line přístup: | https://www.ecice06.com/fileup/1000-3428/PDF/20241129.pdf |
| Tagy: |
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!
|
