Brief safety reminders reduced potentially harmful choices by 19 of 20 AI models tested in a Mount Sinai-led study. Across the experiment, the rate fell from 16.6% without the intervention to 10.1% with it, a reduction of 6.5 percentage points.
Mount Sinai publicized the research on October 8. The peer-reviewed study was published online in Communications Medicine on September 26.
The experiment measured how models responded to prepared clinical instructions. It did not measure injuries, treatment outcomes or the frequency of harmful decisions in an operating hospital.
Researchers tested responses to unsafe instructions
The team evaluated open-weight large language models, systems whose model parameters are available for others to use. A large language model generates responses from patterns learned during training.
Researchers tested 601 cases, comprising 501 synthetic variants derived from 50 templates and 100 cases derived from the MIMIC-IV clinical database. Different instruction framings, safety interventions and repeated runs produced more than 10 million outputs.
Models chose among potentially harmful and safe actions. The reminders encouraged verification or escalation when a request conflicted with safety. The experiment addressed obedience to an unsafe instruction, which can produce a poor decision even when a model has relevant medical knowledge.
The volume of responses reflects repeated testing across combinations of conditions. It is not a study of 10 million patients or separate clinical encounters.
A lower error rate still leaves unsafe answers
The intervention reduced harmful selections but left them at 10.1% overall. Results also differed between the synthetic and database-derived cases, making a single headline rate an incomplete description of performance.
Repeated answers were another concern. Across matched groups of ten runs, 80.4% returned the same response category throughout. Some inconsistent groups switched between safe and potentially harmful categories.
That variability complicates evaluation. An acceptable answer in one run cannot establish that the system will respond safely every time it receives the same request.
The paper describes the work as a proof of concept and explicitly states that it does not estimate clinical risk in deployed systems. Prepared multiple-choice cases cannot reproduce every detail of a clinician’s workflow.
Healthcare buyers need evidence about the actual use
For hospitals and software developers, the findings make instructions part of the system to evaluate. Model capability, surrounding prompts and the route for human intervention all affect how a product behaves.
Our earlier coverage of AI in pharmaceutical development examined a related evaluation problem. A strong result on a prepared test does not by itself establish better decisions throughout a real clinical process.
Safety reminders may be an inexpensive intervention to test, but this study provides no estimate of deployment costs, financial savings or patient benefit. Hospitals still need to assess the specific model, instructions, users and tasks involved before relying on its recommendations.
Cover: Representative hospital equipment and server racks, shown separately. Neither photograph depicts Mount Sinai, the tested models or a study participant. Photographs: Jonathan Borba, Brett Sayles. Source (Pexels License). Source (Pexels License). Cropped and combined by Market Business News.