QR Code (код быстрого отклика)

Benchmarking large language model-based agent systems for clinical decision tasks

Abstract Agentic artificial intelligence (AI) systems, designed to autonomously reason, plan, and invoke tools, have shown promise in healthcare, yet systematic benchmarking of their real-world performance remains limited. In this study, we evaluate two such systems: the open-source OpenManus, built...

Полное описание

Сохранить в:
Библиографические подробности
Главные авторы: Yunsong Liu, Zunamys I. Carrero, Xiaofeng Jiang, Dyke Ferber, Georg Wölflein, Li Zhang, Sanddhya Jayabalan, Tim Lenz, Zhouguang Hui, Jakob Nikolas Kather
Формат: Artigo
Язык:Inglês
Опубликовано: Nature Portfolio 2026-02-01
Серии:npj Digital Medicine
Online-ссылка:https://doi.org/10.1038/s41746-026-02443-6
Метки: Добавить метку
Нет меток, Требуется 1-ая метка записи!