Medical knowledge representation enhancement in large language models through clinical tokens optimization
Abstract During the training of medical large language models (LLMs), conventional tokenizers frequently segment domain-specific medical terms into multiple subword tokens, resulting in suboptimal recognition and representation of specialized vocabulary. As a consequence, the model encounters diffic...
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Artigo |
| Lingua: | Inglês |
| Pubblicazione: |
Nature Portfolio
2026-01-01
|
| Serie: | Scientific Reports |
| Soggetti: | |
| Accesso online: | https://doi.org/10.1038/s41598-026-37438-6 |
| Tags: |
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
