Benchmarking large language model-based agent systems for clinical decision tasks
Abstract Agentic artificial intelligence (AI) systems, designed to autonomously reason, plan, and invoke tools, have shown promise in healthcare, yet systematic benchmarking of their real-world performance remains limited. In this study, we evaluate two such systems: the open-source OpenManus, built...
সংরক্ষণ করুন:
| প্রধান লেখক: | , , , , , , , , , |
|---|---|
| বিন্যাস: | Artigo |
| ভাষা: | Inglês |
| প্রকাশিত: |
Nature Portfolio
2026-02-01
|
| মালা: | npj Digital Medicine |
| অনলাইন ব্যবহার করুন: | https://doi.org/10.1038/s41746-026-02443-6 |
| ট্যাগগুলো: |
কোনো ট্যাগ নেই, প্রথমজন হিসাবে ট্যাগ করুন!
|
