Nature Medicine, Published online: 10 August 2026; doi:10.1038/s41591-026-04574-5
In piloting and deploying a large language model within a large medical center, we learned that benchmark-based evaluations are insufficient for monitoring and evaluating interactions driven by clinicians, and that this requires new methods for monitoring performance.

