QR код

Multi-tier dynamic storage of KV cache for LLM inference under resource-constrained conditions

Abstract The scale of large language models (LLMs) continues to grow in response to increasing demands for intelligent applications. When these large models and their intermediate results, such as key-value (KV) caches, are deployed in resource-constrained environments like edge inference scenarios,...

Повний опис

Збережено в:
Бібліографічні деталі
Автори: Junliang Wang, Jiaqi Hu, Qingping Cao, Yuanrui Zhu, Xiancheng Lin
Формат: Artigo
Мова:Inglês
Опубліковано: Springer 2026-01-01
Серія:Complex & Intelligent Systems
Предмети:
Онлайн доступ:https://doi.org/10.1007/s40747-025-02200-4
Теги: Додати тег
Немає тегів, Будьте першим, хто поставить тег для цього запису!