How good is AI ENT referral judgment compared with a front-line doctor’s? Give the same twelve ear, nose, and throat cases to 100 practicing physicians and to four large language models, and the diagnoses come out roughly even. The models pulled ahead on one narrower task: judging when a patient should be sent to a specialist. That is an interesting result—but it comes from standardized written scenarios, not real clinics, so it is better read as a hypothesis worth testing than as a verdict on who triages better. This piece walks through what the study measured, where its limits are, and what it does and does not mean for anyone typing symptoms into a chatbot.
What the AI ENT Referral Study Found
Researchers built twelve clinical vignettes spanning routine and urgent ENT presentations, then had 100 family medicine and emergency physicians—residents and attendings—work through each one with a diagnosis, a management plan, and a referral decision [Hack S, Empowering front-line physicians with AI: Evaluating large language models in everyday ENT care, 2026]. Four models (Gemini-2.0, ChatGPT-4.0, ChatGPT-5, and OpenEvidence) received identical prompts. A blinded expert panel scored every anonymized response using a structured medical-AI rating tool.
The physicians did well on the parts you would expect. Their mean diagnostic accuracy was 91.6%, and their management accuracy was 87.9% [Hack S, 2026]. Naming the condition and outlining first steps were not the problem.
Referral was the softer spot, and it broke in two directions. In the non-urgent cases, 30.4% of physician responses were over-referrals—sending a patient to a specialist who did not need one [Hack S, 2026]. In the other direction, physicians under-referred when it mattered: in a cerebrospinal fluid (CSF) leak scenario—clear fluid draining from the nose, a possible red flag after head trauma—only about half (52%) recognized the need for urgent escalation [Hack S, 2026]. The models, by the authors’ account, reached comparable diagnostic and management accuracy while reducing both over- and under-referral [Hack S, 2026].
| Dimension | Front-line physicians | Large language models |
|---|---|---|
| Diagnostic accuracy | 91.6% | Comparable (per authors) |
| Management accuracy | 87.9% | Comparable (per authors) |
| Referral appropriateness | Weaker (30.4% over-referral in non-urgent cases; ~48% under-referral in the CSF case) | Higher |
One point before reading too much into that table. The study kept the two kinds of error separate, and they are not equally dangerous. The 30.4% is an over-referral rate, measured across the ten cases that did not need a specialist—costly and inefficient, but rarely dangerous. Under-referral was measured in the two cases that did need urgent escalation, and the CSF leak is the sharp example: roughly half the physicians missed it. The two figures come from different sets of cases, so they are not a head-to-head “48 versus 30″—but the direction of each error is clear, and the under-referral is the one that can hurt a patient.

Why “When to Refer” Is Harder Than Diagnosis
Diagnosis is pattern recognition: match the presentation to the condition. Referral is a second, messier judgment stacked on top of it. It asks how urgent this is, whether it exceeds what primary or emergency care can safely handle, and whether the specialist’s time is warranted now. A doctor making that call is playing two roles at once—patient advocate and system gatekeeper—and those roles pull in opposite directions.
That tension is where errors cluster. Over-refer, and you clog specialist clinics and delay patients who genuinely need the slot. Under-refer, and you miss the CSF leak. A referral rate that looks imperfect is not proof of carelessness; it reflects how genuinely ambiguous these decisions become once the obvious emergencies are set aside.

Clinical Perspective
From a specialist’s standpoint, the referral decision is one of the hardest parts of the encounter to teach, and the study’s pattern is recognizable. Diagnosis has feedback—you eventually learn whether you were right. Referral judgment rarely does: an over-referral looks like caution, and an under-referral often disappears from view until it resurfaces as a complication elsewhere. That missing feedback loop is a plausible reason the skill is slow to build and easy to lose under time pressure.
Why did the models look steadier here? The honest answer is that this study did not establish why. One possible explanation is that a language model produces more standardized, guideline-shaped responses across cases—consistency rather than insight. That is a hypothesis, not a mechanism the study measured, so it should be held loosely.
Two cautions matter more than the headline. First, automation bias: a confidently worded AI recommendation can be over-trusted by clinicians and patients alike, which can entrench an error instead of catching it. Second, these were clean, clinician-style prompts. Results from structured inputs should not be assumed to transfer to a consumer chatbot receiving a partial, emotionally framed symptom description at midnight—a related study by the same group found model performance dropped when prompts were written in informal patient language [Hack S, Evaluating large language models for specialist referral triage in primary care, 2025].
The Limits of This Study
The findings are hypothesis-generating, and several limits keep them there. The cases were twelve standardized vignettes, not real patients, so the true diversity of ENT presentations is compressed. The reference standard for “appropriate” was defined by an expert panel’s rubric rather than an external outcome. “Comparable” diagnostic and management accuracy means similar scores in this sample, not a formal demonstration of statistical equivalence.
Just as important, vignettes strip out the physical examination that real ENT triage leans on—otoscopy, tuning-fork tests, nasal endoscopy, a neck exam, audiometry, and vital signs. Commercial model behavior also shifts by version, date, and region, so a result for one release may not hold for the next. And referral pathways differ sharply by country, insurance system, wait times, and local access to ENT care, which limits how far any single-setting result travels.
What This Means If You’re a Patient
Treat an AI answer as a prompt, not a verdict. It can be a reasonable nudge to ask a clinician, “Should this be seen by an ENT?”—but it cannot examine you, does not know your history, and carries no clinical responsibility. Do not enter identifiable personal health details into a public chatbot.
Some signals should not go to any chatbot at all—they need a person now. Clear fluid draining from the nose after a head injury raises concern for a CSF leak and warrants urgent in-person assessment, often with coordinated ENT, neurosurgery, and trauma input; avoid nose-blowing and straining while you arrange it. Sudden hearing loss in one or both ears, especially within 72 hours, warrants urgent same-day or next-day medical assessment. Rapidly progressive neck swelling, stridor, drooling, a muffled voice, inability to swallow secretions, or difficulty breathing are emergencies—call your local emergency number or go to the emergency department.

Key Takeaways
In a vignette study of twelve ENT cases, large language models and physicians reached comparable diagnostic and management accuracy; the models scored higher on referral appropriateness [Hack S, 2026].
Physicians over-referred in 30.4% of non-urgent cases, and under-referred in the urgent ones—only about half recognized the urgency of a cerebrospinal fluid leak scenario [Hack S, 2026].
Over-referral (30.4%) and under-referral (the CSF miss) are different errors from different cases and are not equally dangerous; the under-referral is the one that can harm a patient. The models reduced both [Hack S, 2026].
The study is hypothesis-generating: standardized written cases, an expert-defined standard, and no physical examination mean it does not show AI is safer in real-world triage.
AI here is decision support, not a replacement for a clinician—and confident AI outputs can invite over-trust.
FAQ
Can AI tell me when to see an ENT specialist? It can offer a reasonable suggestion, and in one vignette study AI referral recommendations scored higher than participating physicians’. But it is a prompt to raise with a clinician, not a decision—AI cannot examine you, weigh your history, or take responsibility for the outcome.
Is ChatGPT accurate for ear, nose, and throat symptoms? On standardized clinical vignettes, models matched physicians on diagnosis and management [Hack S, 2026]. That is a controlled test, not real care; performance also dropped with informal phrasing in a related study, so real-world reliability is less certain than the headline numbers suggest.
Will AI replace doctors for referral decisions? No. The evidence points to support, not substitution. These tools were scored on clean written scenarios; they lack physical examination, real-time history, accountability, and awareness of local referral pathways. The consistent finding is that AI may assist clinical judgment, not replace it.
References
- Hack S, Zalzal HG, Attal R, et al. Empowering front-line physicians with AI: Evaluating large language models in everyday ENT care. Am J Emerg Med. 2026;102:90-97.
- Hack S, Attal R, Yogev D, et al. Evaluating large language models for specialist referral triage in primary care: a quantitative study using otolaryngology scenarios. Fam Pract. 2025;42(6).
- Hack S, Attal R, Steckbeck RJ, et al. The Utility of Large Language Models to Assist With Emergency Triage Decisions Within Otolaryngology. Otolaryngol Head Neck Surg. 2026.
For more interesting content:
https://curiousmd.com/chatgpt-medical-triage-ent-review/
https://curiousmd.com/future-of-clinical-decision-support/
Link out to:
https://med.stanford.edu/news/all-news/2025/02/physician-decision-chatbot.html
https://theconversation.com/dr-chatgpt-is-getting-remarkably-good-at-diagnosing-health-problems-but-actual-doctors-are-still-better-at-weighing-treatment-options-281813
https://www.statnews.com/2025/01/31/chatgpt-beats-doctors-diagnosis-ai-history-medicine-technology/
Joonpyo Hong, MD is a board-certified otolaryngologist practicing in Korea. This article reflects his clinical interpretation of published research and does not constitute individual medical advice.
