Benchmarking Reference-Free LLM Agent Robustness Under Schema, Policy, and Toolset Drift
Tool-using large language model (LLM) agents fail when external interfaces evolve, yet most evaluations emphasize static competence rather than adaptation under drift. We studied this problem with controlled schema, policy, and toolset perturbations derived from tau2-bench retail tasks and distingui...
Uloženo v:
| Hlavní autoři: | , , , |
|---|---|
| Médium: | Artigo |
| Jazyk: | Inglês |
| Vydáno: |
IEEE
2026-01-01
|
| Edice: | IEEE Access |
| Témata: | |
| On-line přístup: | https://ieeexplore.ieee.org/document/11534189/ |
| Tagy: |
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!
|
