PPubMed8 Sep 2026
Performance, Failures, and Oversight of a Large Language Model Agent for Clinical Data Analysis: Evaluation Study
Yilan Wu, Dun Jack Fu, Yukun Zhou, Siegfried K Wagner + 2 more
Background Large language model (LLM) agents capable of generating and executing statistical code from natural language may broaden access to clinical data analysis, yet which pipeline stages they perform reliably and which require expert oversight remain poorly defined. Objective This study aimed to evaluate the…