Inference Optimization for Large Models Based on Adaptive Tensor Swapping and Recomputation
Large Language Models (LLM) have demonstrated outstanding performance in natural language processing tasks. However, their extremely large parameter scales pose a significant challenge because the limited capacity of GPU memory becomes a performance bottleneck for inference tasks. To address this is...
Na minha lista:
| Autor principal: | |
|---|---|
| Formato: | Artigo |
| Idioma: | Inglês |
| Publicado em: |
Editorial Office of Computer Engineering
2025-10-01
|
| coleção: | Jisuanji gongcheng |
| Assuntos: | |
| Acesso em linha: | https://www.ecice06.com/fileup/1000-3428/PDF/jsjgc-51-10-27.pdf |
| Tags: |
Sem tags, seja o primeiro a adicionar uma tag!
|
