Noise-Aware Direct Preference Optimization for RLAIF
Reinforcement Learning from Human Feedback (RLHF) produces powerful instruction-following models but relies on a preference-labeling process that is both costly and slow. An effective alternative, Reinforcement Learning from AI Feedback (RLAIF), uses large language models as teachers for relabeling;...
Guardat en:
| Autors principals: | , , , |
|---|---|
| Format: | Artigo |
| Idioma: | Inglês |
| Publicat: |
MDPI AG
2025-09-01
|
| Col·lecció: | Applied Sciences |
| Matèries: | |
| Accés en línia: | https://www.mdpi.com/2076-3417/15/19/10328 |
| Etiquetes: |
Sense etiquetes, Sigues el primer a etiquetar aquest registre!
|
