QR kód

Benchmarking Reference-Free LLM Agent Robustness Under Schema, Policy, and Toolset Drift

Tool-using large language model (LLM) agents fail when external interfaces evolve, yet most evaluations emphasize static competence rather than adaptation under drift. We studied this problem with controlled schema, policy, and toolset perturbations derived from tau2-bench retail tasks and distingui...

Celý popis

Uloženo v:
Podrobná bibliografie
Hlavní autoři: Mohammad Hasbi Assidiqi, Daniyal Alghazzawi, Suaad Alarifi, Li Cheng
Médium: Artigo
Jazyk:Inglês
Vydáno: IEEE 2026-01-01
Edice:IEEE Access
Témata:
On-line přístup:https://ieeexplore.ieee.org/document/11534189/
Tagy: Přidat tag
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!