←  Back

Publications

The work behind
the benchmark

Peer-reviewed research from our team on what clinical AI actually costs, what it orders, and where accuracy stops being the right question.

NEJM AI · In press

Comparative Cost of Care Recommended by Large Language Models vs. Treating Physicians

Dvir Aran, Jeremy D. Goldhaber-Fiebert, Ittai Many, Anthony Nguyen, Daniel Yang, Shahar Shelly

Twenty-four AI systems wrote plans for 200 real outpatient and emergency visits, priced at 2026 Medicare rates. Their care cost $314 per visit against $97 for the treating physician — 3.2× — driven mostly by testing at the 66% of visits where the physician ordered none. Cost-aware prompting narrowed the gap but never closed it.

Link on publication

Nature Medicine · Published

Prospective evaluation of a large language model clinical decision support system in the emergency department

Liron Leibovitch, Adi Ahituv, Alon Gorenshtein, Dvir Aran, Moran Sorka, Keren Miron, Shahar Shelly

A DECIDE-AI stage 1 evaluation across 1,138 emergency department patients over four weeks. Outputs were rated clinically appropriate in 99 of 100 sampled cases — yet clinician adoption fell from 68% to 30%. Sustained engagement, not algorithmic accuracy, was the binding constraint.

Read the paper →

Communications Medicine · Published

Clinical liability cases reveal a coupling between legal defensibility and medical procedure escalation in large language models

Dvir Aran, Ronen Perry, Shahar Shelly

Fifteen models, 198 US malpractice cases, 3,072 simulated consultations. The models that best matched court-determined standards of care were also the models that ordered the most — 9.3 procedures at $1,118 versus 1.3 at $221 (ρ = 0.95). The coupling tightens with each model generation, and replicates on UK cases.

Read the paper →

J. Am. Med. Inform. Assoc. · Published

DiagnosticXchange: an open-source framework for evaluating safety, efficiency, and diagnostic reasoning in clinical AI systems

Moran Sorka, Alon Gorenshtein, Hillel Abramovitch, Pannathat Soontrapa, Shahar Shelly, Dvir Aran

A dynamic simulation in which AI systems work up 216 peer-reviewed cases by ordering real tests, imaging and procedures, each priced by CPT code. Three systems indistinguishable on accuracy (93.5–94.0%) differed 1.75-fold in cost and 2.1-fold in physician effort — differences no accuracy benchmark can see.

Read the paper →

PLOS Digital Health · Published

A multi-agent approach to neurological clinical reasoning

Moran Sorka, Alon Gorenshtein, Dvir Aran, Shahar Shelly

Ten models against 305 neurology board certification questions, each graded on three axes of complexity. Splitting the reasoning across specialised agents — analysis, retrieval, synthesis, validation — carried LLaMA 3.3-70B from 69.5% to 89.2%, and evened out the subspecialty gaps that retrieval alone never closed.

Read the paper →

JMIR AI · Published

GPT-4 as a Clinical Decision Support Tool in Ischemic Stroke Management: Evaluation Study

Amit Haim Shmilovitch, Mark Katson, Michal Cohen-Shelly, Shlomi Peretz, Dvir Aran, Shahar Shelly

One hundred acute stroke presentations, each read by GPT-4 and scored against a stroke specialist and the care actually delivered. Agreement on thrombectomy reached AUC 0.94, and its 90-day mortality ranking (AUC 0.89) outperformed the supervised risk models in clinical use.

Read the paper →