HybridHiT-UNet: Multi-Scale Temporal U-Net with Hierarchical Shot-Aware Transformers for Video Summarization
Video summarization aims to produce a short yet informative summary of a long video while reducing the amount of redundancy. Most transformer-based methods are single-temporal scale or are unconcerned with shot-level structure, limiting temporal coherence and cross-dataset generalization. To fill th...
Shranjeno v:
| Principais autores: | , , , |
|---|---|
| Format: | Artigo |
| Jezik: | Inglês |
| Izdano: |
MDPI AG
2026-05-01
|
| Serija: | Machine Learning and Knowledge Extraction |
| Teme: | |
| Online dostop: | https://www.mdpi.com/2504-4990/8/5/135 |
| Oznake: |
Brez oznak, prvi označite!
|
