An agent framework called Asclepius improved the completion of critical clinical actions in simulated emergency-department shifts, where models must manage several patients under continuous time and resource pressure.
The researchers found that current agents often reached the correct diagnosis but failed to complete every necessary action on time. They tracked three long-horizon problems: drifting away from instructions, incomplete treatment and differences in response time associated with case severity.
Asclepius rewrites its operating manual between shifts using feedback from previous traces, keeps high-stakes regimen knowledge in an external clinical skills library and divides per-turn decisions among three isolated subagents across the patient queue. All three components were needed for the clearest reduction in the coupled failure modes.
On held-out batches not used during harness evolution, critical-action correctness rose by 22% over a strong baseline while diagnostic accuracy was preserved; the reported p-value was 0.024. Across all 10 batches, critical actions improved by 25% and timeliness by 13%, with gains judged across five models from three families. These are simulation results from a new arXiv preprint, not evidence for autonomous clinical deployment. Real care would require prospective safety validation and accountable human oversight.