Hardware-Friendly Fully Quantized Mamba-2 Model and Its FPGA-Based Accelerator
In this paper, we propose a fully quantized Mamba-2 model to enable efficient large language model (LLM) processing for edge AI. As a promising next-generation architecture, Mamba offers a scalable alternative to Transformer-based models with better suitability for edge AI. In the proposed model, th...
Sábháilte in:
| Príomhchruthaitheoirí: | , |
|---|---|
| Formáid: | Artigo |
| Teanga: | Inglês |
| Foilsithe / Cruthaithe: |
IEEE
2026-01-01
|
| Sraith: | IEEE Access |
| Ábhair: | |
| Rochtain ar líne: | https://ieeexplore.ieee.org/document/11534219/ |
| Clibeanna: |
Níl clibeanna ann, Bí ar an gcéad duine le clib a chur leis an taifead seo!
|
