Summon a demon and bind it: A grounded theory of LLM red teaming.
Engaging in the deliberate generation of abnormal outputs from Large Language Models (LLMs) by attacking them is a novel human activity. This paper presents a thorough exposition of how and why people perform such attacks, defining LLM red-teaming based on extensive and diverse evidence. Using a for...
שמור ב:
| Principais autores: | , , |
|---|---|
| פורמט: | Artigo |
| שפה: | Inglês |
| יצא לאור: |
Public Library of Science (PLoS)
2025-01-01
|
| סדרה: | PLoS ONE |
| גישה מקוונת: | https://escholarship.org/uc/item/2dh5j3rn |
| תגים: |
אין תגיות, היה/י הראשונ/ה לתייג את הרשומה!
|
