Benchmarking large language model-based agent systems for clinical decision tasks
Abstract Agentic artificial intelligence (AI) systems, designed to autonomously reason, plan, and invoke tools, have shown promise in healthcare, yet systematic benchmarking of their real-world performance remains limited. In this study, we evaluate two such systems: the open-source OpenManus, built...
Сохранить в:
| Главные авторы: | , , , , , , , , , |
|---|---|
| Формат: | Artigo |
| Язык: | Inglês |
| Опубликовано: |
Nature Portfolio
2026-02-01
|
| Серии: | npj Digital Medicine |
| Online-ссылка: | https://doi.org/10.1038/s41746-026-02443-6 |
| Метки: |
Нет меток, Требуется 1-ая метка записи!
|
