QRコード

Benchmarking large language model-based agent systems for clinical decision tasks

Abstract Agentic artificial intelligence (AI) systems, designed to autonomously reason, plan, and invoke tools, have shown promise in healthcare, yet systematic benchmarking of their real-world performance remains limited. In this study, we evaluate two such systems: the open-source OpenManus, built...

詳細記述

保存先:
書誌詳細
主要な著者: Yunsong Liu, Zunamys I. Carrero, Xiaofeng Jiang, Dyke Ferber, Georg Wölflein, Li Zhang, Sanddhya Jayabalan, Tim Lenz, Zhouguang Hui, Jakob Nikolas Kather
フォーマット: Artigo
言語:Inglês
出版事項: Nature Portfolio 2026-02-01
シリーズ:npj Digital Medicine
オンライン・アクセス:https://doi.org/10.1038/s41746-026-02443-6
タグ: タグ追加
タグなし, このレコードへの初めてのタグを付けませんか!