Autocorrelation Matrix Knowledge Distillation: A Task-Specific Distillation Method for BERT Models
Pre-trained language models perform well in various natural language processing tasks. However, their large number of parameters poses significant challenges for edge devices with limited resources, greatly limiting their application in practical deployment. This paper introduces a simple and effici...
Gorde:
| Egile Nagusiak: | , , , |
|---|---|
| Formatua: | Artigo |
| Hizkuntza: | Inglês |
| Argitaratua: |
MDPI AG
2024-10-01
|
| Saila: | Applied Sciences |
| Gaiak: | |
| Sarrera elektronikoa: | https://www.mdpi.com/2076-3417/14/20/9180 |
| Etiketak: |
Etiketarik gabe, Izan zaitez lehena erregistro honi etiketa jartzen!
|
