Fast collaborative inference via distributed speculative decoding
Speculative decoding accelerates Large Language Model (LLM) inference by allowing a lightweight draft model to predict multiple future tokens that are subsequently verified by a larger target model. In AI-native Radio Access Networks (AI-RAN), this mechanism naturally enables device-edge collaborati...
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Artigo |
| Sprache: | Inglês |
| Veröffentlicht: |
KeAi Communications Co., Ltd.
2026-01-01
|
| Schriftenreihe: | Journal of Information and Intelligence |
| Schlagworte: | |
| Online-Zugang: | http://www.sciencedirect.com/science/article/pii/S2949715925000782 |
| Tags: |
Keine Tags, Fügen Sie das erste Tag hinzu!
|
