Image credit: Denis Pobytov / DigitalVision Vectors / Getty
Around the world, technology companies are racing to develop artificial intelligence (AI)-enabled tools for clinical decision support (CDS) for a global market that is forecast to reach US $15 billion by 2033. A new generation of AI tools powered by large language models (LLMs) is moving rapidly from development into clinical settings. Health systems and technology companies are integrating these tools into clinical workflows, promising faster evidence synthesis and more-responsive decision-making at the bedside. But as CDS systems enter into practice, fundamental questions remain about how they are evaluated, what evidence supports their use, and how they will be governed at scale.
Among the US leaders vying for market share are generalist consumer AI tools such as Chat GPT, as well as specialized CDS systems, including those from Abridge, Atropos Health and OpenEvidence. These specialist CDS companies have diverging opinions on how the field will develop next.
Pittsburgh-based Abridge develops ambient AI technology that listens to clinician–patient conversations and automatically generates clinical notes and other workflow-support tools for health systems. Boston-based OpenEvidence markets an AI tool directly to physicians that the company claims is the “most widely used medical AI” among US clinicians and is focused on distilling the best available evidence in the published literature. Atropos Health, a company spun out of Stanford University, promotes a CDS system designed to enable clinicians to ask novel questions in situations in which gaps may exist in the traditional, formally published evidence.
The data that power CDS systems increasingly come directly from patient records. Over the past decade, electronic health records (EHR) that automatically scan patient data to track disease indications and issue alerts to clinicians have become ubiquitous in many health systems worldwide. But it was not until about 3 years ago, with the mass adoption of ‘ambient scribes’ that automatically transcribe clinician–patient consultations into clinical records, that AI-powered CDS seemed poised for prime time. “Looking ahead,” says Daniel Nadler, CEO and founder of OpenEvidence, his company’s goal is to “automate much of the ‘standard of care’ parts of evidence-based practice and bring physicians back to the most human parts of medicine.”
Entering the clinical workflow
The fate of CDS tools ultimately lies in the hands of clinicians, emphasizes Kevin Johnson, a professor of pediatrics and biomedical informatics at the University of Pennsylvania who has probed CDS systems in numerous studies. Just because a CDS app is downloaded by clinicians does not mean it is trusted and useful, he warns.
Citing a 2024 study that found physicians override approximately 90% of drug–drug interaction alerts generated by CDS tools, Johnson says clinicians routinely disregard CDS systems — whether their decision support is based on AI or other software methods — because they often inaccurately prioritize information, especially when dealing with highly specific medical questions. “The system will tell me that this patient maybe would be better with this med than another med,” he explains, “and somewhere in the neighborhood of 80 to 90 percent of the time, the clinician overrides that.”
Trust issues featured prominently in a May 2026 audit of the rollout of 20 government-approved AI scribes among 30,000 public health doctors in Ontario, Canada, serving more than 16 million patients. The audit found that all 20 approved AI scribe vendors had inaccuracies in their system-generated notes, including hallucinated content, incorrect treatment suggestions and missed clinical details that could impact patients’ health. None of the systems were sufficiently evaluated by Ontario health officials to ensure they mitigated the risk of creating unfair or biased outcomes.
In California, a class action lawsuit filed in April 2026 against three California health systems has challenged the use of AI scribes by providers without patient knowledge or explicit consent. Karandeep Singh, a practicing nephrologist and chief health AI officer at UC San Diego Health, worries that understudied AI CDS systems are being adopted in a haphazard fashion that raises concerns about patients’ safety. “Clinicians are using these tools in an unregulated sort of way,” says Singh, who wrote a March 2026 editorial in which he noted that many clinicians are using general-purpose consumer AI tools without any institutional oversight.
Historically, says Singh, health systems have looked at AI, and generative AI especially, as risky, and they have implemented it cautiously. “What they are now faced with is the conundrum that clinicians are not waiting,” says Singh.
What clinicians are using
Although their products are not available in Europe, OpenEvidence and Abridge have both forged relationships with leading medical publishers, including the New England Journal of Medicine and JAMA. More recently, OpenEvidence also announced a partnership with Springer Nature (the publisher of Nature Medicine). “The thought here is, let’s take this mountain of data that’s available to you and let’s shape it to the patient [who’s] in the room,” says Abridge’s clinical strategy director Matt Troup.
In sharp contrast, Atropos’ CEO and co-founder Brigham Hyde argues that there is a paucity of real-world patient data underlying many AI CDS systems, even when they are linked with major medical publishers.
To address this issue, the Atropos engineering team is assembling a novel medical library it calls Alexandria, which contains reams of unpublished observational evidence on a vast array of topics. These retrospective observations are generated in response to clinical questions, using de-identified patient data from Stanford Health Care. Each evidence summary is reviewed by AI. Like the Abridge CDS, the Atropos CDS is marketed to institutional entities such as health systems and hospitals rather than directly to clinicians.
Atropos is confident their approach works best when answers are not in the published literature. In a May 2025 study funded by Atropos, researchers at hospitals in Ann Arbor, New York City, San Diego and Toronto worked with emergency medicine, pathology and clinical informatics specialists at Stanford Health Care (some employed by Atropos) to compare how five types of LLMs responded to a set of 50 clinical questions. The study reported that while OpenEvidence outperformed three general-purpose LLMs (ChatGPT-4, Claude 3 Opus and Gemini 1.5 Pro), as well as Atropos, in situations in which answers could be found in the published literature, in situations in which published data were not available, Atropos delivered the most medically reliable answers. Atropos is now working with partners such as Meta and Microsoft to integrate the Alexandria library into clinical workflows.
In a recent Nature Medicine study, three frontier LLMs (GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6) were compared with OpenEvidence and another specialist tool, UpToDate Expert AI from Wolters Klewer. The authors reported that the frontier models outperformed the specialist tools in tests that relied on three public benchmarks. The study did not assess citation quality, and two of the benchmarks may have been seen by the LLMs during training. However, the study raises questions about how best to evaluate CDS tools. In a closed-to-comments LinkedIn post, OpenEvidence alleges the paper’s authors have a “massive undisclosed conflict of interest” and that the paper contains methodological flaws.
The data challenge
Jackie Gerhart, the chief medical officer at Epic, which markets an EHR system that manages health information from about 300 million patients, watches the competition between AI CDS innovators with keen interest. Much of what is currently hyped as ‘new’ in CDS innovation is not particularly new, she says. “When you’re thinking about companies like OpenEvidence or UpToDate or curated knowledge bases that can be looked at through an AI lens,” she reflects, “we view the companies that do that more as sort of like content-forwarding companies. They’re curating content from journals.”
Harvesting and harnessing patient data directly from EHRs, as is being done at Stanford Health Care, could substantially advance CDS innovation, she agrees. But she cautions about the need for great care when highly confidential, tightly regulated patient records are used for any sort of CDS innovation. “There [are] rules of the road about where and when that information can be shared,” she says. “But one place it can be shared is within the actual workflow of a patient that might have thousands of patients that are very similar to them, trying to see what their trajectory has been.”
