Código QR (código de barras bidimensional)

Inference Optimization for Large Models Based on Adaptive Tensor Swapping and Recomputation

Large Language Models (LLM) have demonstrated outstanding performance in natural language processing tasks. However, their extremely large parameter scales pose a significant challenge because the limited capacity of GPU memory becomes a performance bottleneck for inference tasks. To address this is...

ver descrição completa

Na minha lista:
Detalhes bibliográficos
Autor principal: LIANG Xuning, WANG Siqi, YANG Hailong, LUAN Zhongzhi, LIU Yi, QIAN Depei
Formato: Artigo
Idioma:Inglês
Publicado em: Editorial Office of Computer Engineering 2025-10-01
coleção:Jisuanji gongcheng
Assuntos:
Acesso em linha:https://www.ecice06.com/fileup/1000-3428/PDF/jsjgc-51-10-27.pdf
Tags: Adicionar Tag
Sem tags, seja o primeiro a adicionar uma tag!