Multi-tier dynamic storage of KV cache for LLM inference under resource-constrained conditions
Abstract The scale of large language models (LLMs) continues to grow in response to increasing demands for intelligent applications. When these large models and their intermediate results, such as key-value (KV) caches, are deployed in resource-constrained environments like edge inference scenarios,...
Збережено в:
| Автори: | , , , , |
|---|---|
| Формат: | Artigo |
| Мова: | Inglês |
| Опубліковано: |
Springer
2026-01-01
|
| Серія: | Complex & Intelligent Systems |
| Предмети: | |
| Онлайн доступ: | https://doi.org/10.1007/s40747-025-02200-4 |
| Теги: |
Немає тегів, Будьте першим, хто поставить тег для цього запису!
|
