Opinion: Autonomous AI will beat AI-assisted physicians at some medical tasks by 2030

Opinion: Autonomous AI will beat AI-assisted physicians at some medical tasks by 2030

Ezekiel J. Emanuel and Abe Butler-Baker have been debating the proper place for AI in medicine with American Medical Association CEO John Whyte. Now, they are taking their discussion to STAT’s First Opinion. Read Emanuel and Butler’s essay below and read Whyte’s essay here. In 1867, Joseph Lister published his research on carbolic acid and antiseptic surgical technique. In September 1871, he was summoned to Queen Victoria, who had a rapidly growing abscess in her left armpit. Using his antiseptic surgical technique, Joseph Lister successfully drained the pus. Queen Victoria recovered without fever or other complications. The antiseptic technique quickly gained approval in the U.K. and Europe, but not among American physicians. In fall 1876, Lister arrived in the U.S. for an international medical meeting and traveled to Philadelphia, New York, and Boston where he demonstrated his antiseptic technique to leading American physicians. In fall 1881, 14 years after Lister’s publication and five years after his American demonstrations, President Garfield was shot. The nation’s leading surgeons of the day attended him. Skeptical of Lister’s innovation despite the strong data supporting it, these haughty surgeons repeatedly stuck their unwashed hands and unsterilized instruments into Garfield’s wound to try to locate the bullet. With good intentions but ignoring science, they infected President Garfield and killed him. In August, we co-authored an article in JAMA with Vinod Khosla and Neal Khosla arguing that autonomous artificial intelligence will likely soon exceed both unaided physicians and physicians aided by AI at five core cognitive medical tasks: patient information gathering, differential diagnosis, selecting cost-efficient tests, prescribing guideline-concordant treatment, and managing chronic illnesses. We reached this conclusion by evaluating all studies comparing autonomous AI to physicians with or without AI published since January 2024. Based on the data, we project autonomous AI will be superior to AI-aided physicians at, and ready for real-world implementation, on some, maybe all, of these five tasks by 2030. The preponderance of studies shows autonomous AI beating unaided physicians at the five cognitive medical tasks. Data from medicine and other fields including chess suggest that once autonomous AI beats unaided humans, autonomous AI shortly thereafter consistently beats human-AI hybrids. While regulatory barriers and physician resistance have limited real-world trials of autonomous AI, the data we do have is remarkable. A few examples: Google’s AMIE is statistically significantly better than physicians at eliciting all portions of patient history. ChatGPT beats physicians at differential diagnosis by 18 percentage points (92% vs. 74%). Microsoft’s AI Diagnostic Orchestrator produced correct final diagnoses 4.02 times more frequently than physicians (80.4% vs. 20%) at an average testing cost ($2,397) which was 19.1% lower. MIRA, an autonomous AI, prescribed guideline-concordant treatment 35 percentage points more frequently than doctors. And a study at Stanford showed autonomous AI trounced physicians at managing diabetes, achieving a stable insulin dose in 15 days while doctors still could not after eight weeks. And, nine of 13 studies since January 2024 comparing autonomous AI to AI-aided physicians show autonomous AI is superior. For instance, autonomous AI beats physicians with AI at diagnosis by 21.3 percentage points. It sounds counterintuitive — why would having the doctor involved be worse? However, data show that when AI is highly capable, doctors add more errors and false corrections than beneficial insights. AI is rapidly improving, which will intensify this phenomenon. American Medical Association CEO John Whyte, who is a friend, and others have criticized our conclusion that autonomous AI will beat AI-aided physicians. They have three main arguments. First, Whyte and others argue that AI should not practice autonomously because we do not have adequate licensure and liability structures for autonomous AI. We agree. Current licensure and liability structures for autonomous AI are inadequate, but the answer is to create adequate structures, not to reject autonomous AI. Indeed, we proposed such a licensing structure in JAMA and currently have a paper under review proposing a feasible, comprehensive liability structure for autonomous AI. Second, Whyte and colleagues argue doctors will still be needed for the “art of medicine”: trust, emotional awareness, compassion, patient relationships, and communicating difficult decisions. As a result, they say, doctors will still need to be “on-the-loop.” Studies indicate that AI can handle difficult conversations and demonstrate empathy better than doctors. In 2025, Alastair Howcroft and colleagues found that 13 out of 15 studies comparing empathy in medicine showed statistically significantly higher empathy ratings for AI than human health care professionals. Patient actors felt more at ease (97% vs. 65%, p < 0.001) and listened to (95% vs. 72%, p < 0.001) by Google’s AMIE than by primary care physicians. Humans may be needed for tasks that only humans can perform, like holding patients’ hands, but doctors may not be best for compassion and empathy. Third, Whyte and others dismiss our conclusion because many of the studies examining AI in medicine are simulations. They argue that AI would fail to perform in a real-world clinical environment. While they correctly observe that much of the medical AI literature is simulated, their conclusion is inaccurate. The lack of real-world autonomous AI testing is attributable not to poor capability but rather to doctors, regulators, and laws that restrict such tests. Indeed, the real-world data we do have show autonomous AI beating physicians using AI. In the Annals of Internal Medicine, a team led by Dan Zeltzer studied 461 real patient visits. AI had a structured online chat with the patient and then generated treatment recommendations. Physicians then saw the patient and had access to the AI treatment recommendations, and they generated their own recommendations. Physicians’ treatment recommendations were worse than AI’s, even though physicians had access to the AI’s recommendations. These were real patient cases with information gathered directly from patients by AI. The answer to the “simulated cases” objection is to test autonomous AI in the real world, not to dismiss autonomous AI’s capability. AI will not take over all roles from doctors. After all, AI cannot independently perform surgeries, deliver babies, or conduct physical procedures since it lacks physical capability. But data indicate that autonomous AI will be better than physicians working with AI at core cognitive medical tasks. We need to test autonomous AI at real-world deployment. If we do not, we may harm our patients by denying them what science shows is likely the best medical care, just as American surgeons did to President Garfield in 1881. Ezekiel J. Emanuel is vice provost for global initiatives and co-director of the Healthcare Transformation Institute at the Perelman School of Medicine and the Wharton School at the University of Pennsylvania. Abe Baker-Butler is a research fellow in the Department of Medical Ethics and Health Policy at the University of Pennsylvania.

Original Source

Read the full article at Statnews →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.