QR-koda

FPGA Acceleration With Hessian-Based Comprehensive Intra-Layer Mixed-Precision Quantization for Transformer Models

Recent advancements in using FPGAs as co-processors for language model acceleration, particularly for energy efficiency and flexibility, face challenges due to limited memory capacity. This limitation hinders the deployment of transformer-based language models. To address this challenge, we propose...

Olles dieđut

Furkejuvvon:
Bibliográfalaš dieđut
Váldodahkkit: Woohong Byun, Jongseok Woo, Saibal Mukhopadhyay
Materiálatiipa: Artigo
Giella:Inglês
Almmustuhtton: IEEE 2025-01-01
Ráidu:IEEE Access
Fáttát:
Liŋkkat:https://ieeexplore.ieee.org/document/10973048/
Fáddágilkorat: Lasit fáddágilkoriid
Eai fáddágilkorat, Lasit vuosttaš fáddágilkora!