Código QR

Benchmarking Reference-Free LLM Agent Robustness Under Schema, Policy, and Toolset Drift

Tool-using large language model (LLM) agents fail when external interfaces evolve, yet most evaluations emphasize static competence rather than adaptation under drift. We studied this problem with controlled schema, policy, and toolset perturbations derived from tau2-bench retail tasks and distingui...

ver descrição completa

Na minha lista:
Detalhes bibliográficos
Principais autores: Mohammad Hasbi Assidiqi, Daniyal Alghazzawi, Suaad Alarifi, Li Cheng
Formato: Artigo
Idioma:Inglês
Publicado em: IEEE 2026-01-01
Colecção:IEEE Access
Assuntos:
Acesso em linha:https://ieeexplore.ieee.org/document/11534189/
Tags: Adicionar Tag
Sem tags, seja o primeiro a adicionar uma tag!