QR-Code

Fast collaborative inference via distributed speculative decoding

Speculative decoding accelerates Large Language Model (LLM) inference by allowing a lightweight draft model to predict multiple future tokens that are subsequently verified by a larger target model. In AI-native Radio Access Networks (AI-RAN), this mechanism naturally enables device-edge collaborati...

Ausführliche Beschreibung

Gespeichert in:
Bibliografische Detailangaben
Hauptverfasser: Ce Zheng, Ke Zhang, Chen Sun, Wenqi Zhang, Qiong Liu, Angesom Ataklity Tesfay
Format: Artigo
Sprache:Inglês
Veröffentlicht: KeAi Communications Co., Ltd. 2026-01-01
Schriftenreihe:Journal of Information and Intelligence
Schlagworte:
Online-Zugang:http://www.sciencedirect.com/science/article/pii/S2949715925000782
Tags: Tag hinzufügen
Keine Tags, Fügen Sie das erste Tag hinzu!