Summon a demon and bind it: A grounded theory of LLM red teaming.
Engaging in the deliberate generation of abnormal outputs from Large Language Models (LLMs) by attacking them is a novel human activity. This paper presents a thorough exposition of how and why people perform such attacks, defining LLM red-teaming based on extensive and diverse evidence. Using a for...
Na minha lista:
| Principais autores: | , , |
|---|---|
| Formato: | Artigo |
| Idioma: | Inglês |
| Publicado em: |
Public Library of Science (PLoS)
2025-01-01
|
| Colecção: | PLoS ONE |
| Acesso em linha: | https://escholarship.org/uc/item/2dh5j3rn |
| Tags: |
Sem tags, seja o primeiro a adicionar uma tag!
|
