Benchmarking Reference-Free LLM Agent Robustness Under Schema, Policy, and Toolset Drift
Tool-using large language model (LLM) agents fail when external interfaces evolve, yet most evaluations emphasize static competence rather than adaptation under drift. We studied this problem with controlled schema, policy, and toolset perturbations derived from tau2-bench retail tasks and distingui...
Na minha lista:
| Principais autores: | , , , |
|---|---|
| Formato: | Artigo |
| Idioma: | Inglês |
| Publicado em: |
IEEE
2026-01-01
|
| Colecção: | IEEE Access |
| Assuntos: | |
| Acesso em linha: | https://ieeexplore.ieee.org/document/11534189/ |
| Tags: |
Sem tags, seja o primeiro a adicionar uma tag!
|
