Benchmarking large language model-based agent systems for clinical decision tasks
Abstract Agentic artificial intelligence (AI) systems, designed to autonomously reason, plan, and invoke tools, have shown promise in healthcare, yet systematic benchmarking of their real-world performance remains limited. In this study, we evaluate two such systems: the open-source OpenManus, built...
保存先:
| 主要な著者: | , , , , , , , , , |
|---|---|
| フォーマット: | Artigo |
| 言語: | Inglês |
| 出版事項: |
Nature Portfolio
2026-02-01
|
| シリーズ: | npj Digital Medicine |
| オンライン・アクセス: | https://doi.org/10.1038/s41746-026-02443-6 |
| タグ: |
タグなし, このレコードへの初めてのタグを付けませんか!
|
