QubitCache: Quantum-Inspired Probabilistic Attention Preservation for KV-Cache Compression
Large language model inference faces a critical memory bottleneck from KV cache, which grows linearly with sequence length and dominates GPU memory during long-context generation. Existing compression methods reduce memory through token eviction but irreversibly discard attention relationships essen...
-д хадгалсан:
| Үндсэн зохиолчид: | , , , |
|---|---|
| Формат: | Artigo |
| Хэл сонгох: | Inglês |
| Хэвлэсэн: |
IEEE
2026-01-01
|
| Цуврал: | IEEE Access |
| Нөхцлүүд: | |
| Онлайн хандалт: | https://ieeexplore.ieee.org/document/11479587/ |
| Шошгууд: |
Шошго байхгүй, Энэхүү баримтыг шошголох эхний хүн болох!
|
