AB jailbreaking - a novel hybrid framework for exploitation of adversarial vulnerabilities in LLMs
Abstract Large language models (LLMs) have advanced rapidly but remain vulnerable to adversarial “jailbreaking” attacks that elicit harmful or disallowed outputs. We propose AB-JB, a three-stage hybrid jailbreak framework that combines black-box semantic adversarial prompt variant generation with a...
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Artigo |
| Lenguaje: | Inglês |
| Publicado: |
Nature Portfolio
2026-04-01
|
| Colección: | Scientific Reports |
| Materias: | |
| Acceso en línea: | https://doi.org/10.1038/s41598-026-44403-w |
| Etiquetas: |
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
