QR Kod

FPGA Acceleration With Hessian-Based Comprehensive Intra-Layer Mixed-Precision Quantization for Transformer Models

Recent advancements in using FPGAs as co-processors for language model acceleration, particularly for energy efficiency and flexibility, face challenges due to limited memory capacity. This limitation hinders the deployment of transformer-based language models. To address this challenge, we propose...

Ful tanımlama

Kaydedildi:
Detaylı Bibliyografya
Asıl Yazarlar: Woohong Byun, Jongseok Woo, Saibal Mukhopadhyay
Materyal Türü: Artigo
Dil:Inglês
Baskı/Yayın Bilgisi: IEEE 2025-01-01
Seri Bilgileri:IEEE Access
Konular:
Online Erişim:https://ieeexplore.ieee.org/document/10973048/
Etiketler: Etiketle
Etiket eklenmemiş, İlk siz ekleyin!