AI-generated laryngoscopy images are now realistic enough that a group of 114 clinicians struggled to sort them from the real thing. Shown a mix of authentic videostroboscopic frames of the larynx and synthetic ones produced by a generative model, the clinicians identified the authentic frames 70.9% of the time. On the synthetic frames, accuracy fell to 57.8% ([Morrison, Human Evaluation of Synthetic Videostroboscopic Laryngeal Images Generated Using StyleGAN3, 2026]).
The two numbers say different things. Clinicians held onto a reasonable recognition of authentic tissue; it was the synthetic images that slipped through a much wider gate.
AI-Generated Laryngoscopy Images: What the Study Actually Measured
A team at Weill Cornell’s Sean Parker Institute for the Voice, working with computational collaborators, trained StyleGAN3, a generative adversarial network, on two age-stratified sets of laryngeal images: one drawn from patients aged 65 and over, the other from patients under 65. From a pool of 144 images containing both authentic frames and generated ones sampled at five training checkpoints, each participant received a randomized, balanced 36-image survey.
Overall accuracy across the whole survey came to 59.9%, with a standard deviation of 14.2 points. The gap between real and synthetic identification was statistically significant (p<0.001). No meaningful difference emerged between the two age-stratified datasets (p=0.27), though a null result of this kind means a difference was not detected, not that none exists.
One boundary matters for everything downstream: this study measured human perception, not diagnosis. Nobody was asked whether a lesion was benign or malignant. The published level of evidence is listed as not applicable. The question was whether an image looks real, which is narrower and more specific than it first appears.
More Training Did Not Mean More Realism
The intuitive assumption about generative models is that longer training yields better output. The data here point elsewhere.
At the earliest checkpoint (5,120 kimg), clinicians correctly flagged synthetic images 79.6% of the time, since the model was still producing obvious tells. By 10,120 kimg that had fallen to 59.1%, a significant drop (p<0.001), and that is where the steep part of the curve ended. Realism continued to edge upward to its peak at 20,120 kimg, where clinician accuracy sat at 44.8%, below the 50% mark expected from pure guessing, but each additional round of training was buying progressively less.

The pattern was consistent within this experiment, though it remains a single study. Its practical implication is uncomfortable: producing images convincing enough to fool experienced observers in a survey setting does not require an enormous training budget, because the realism plateau arrives early.

Accuracy Differed by Device
The subgroup analysis turned up something easy to overlook. Participants who took the survey on a computer were correct 63.9% of the time. Those on a phone, 56.2% (p=0.014). A confounder analysis indicated that device choice and specialty were statistically independent (p=0.164), so the phone effect was not simply a proxy for which specialty the respondent came from.
The caveat is real: device was self-selected, not randomized. Someone answering on a phone between cases may be reading differently than someone at a workstation, and screen size may not be the operative variable at all. The direction of the effect, however, is consistent.

The Radiology Deepfake Study, and Why It Is Not the Same Ruler
Five months before this paper appeared, Radiology published a study in which 17 radiologists from six countries evaluated synthetic radiographs ([Tordjman, The Rise of Deepfake Medical Imaging, 2026]). It drew wide coverage, some of it under headlines suggesting radiologists could not tell the difference at all. The paper says something more measured. Once informed that synthetic images were present, the radiologists reached 75% accuracy on one dataset (95% CI: 68, 81) and 70% on another (95% CI: 62, 78), with no evidence of a difference between them (p=.07). And 41% of them, seven of 17, had flagged the presence of AI-generated images spontaneously, while still blinded to the study’s purpose.
These two studies cannot be placed on the same ruler, because they used different generative architectures, reader populations, task structures, and accuracy calculations. Setting 59.9% beside 75% would be a category error. What can be said is narrower and still meaningful: the same question is now being asked independently in two imaging specialties, and neither has produced a reassuring answer.
Clinical Perspective: What a Still Frame Leaves Out
One feature of the study design limits how far its findings should be carried.
The participants judged still frames. Stroboscopy is not a still-frame modality. It exists to visualize vocal fold vibration during phonation: the mucosal wave, phase symmetry, and the pattern of glottal closure across the cycle ([Mehta, Current role of stroboscopy in laryngeal imaging, 2012]). What a stroboscopy examination assesses is motion, and a single frame carries the anatomy while discarding the physiology.
That limits the claim in both directions. On the forgery side, a convincing still frame is not a convincing stroboscopy examination, since the modality’s diagnostic value lies in the vibratory features a single frame cannot contain. On the enthusiasm side, if a generative model has learned to reproduce laryngeal appearance without reproducing laryngeal behavior, its value as training data for any system meant to assess vibration is an open question rather than a settled benefit.
Both of those inferences are a clinical reading of what the study’s design allows, not conclusions the study itself draws. The authors tested perception; the bounds proposed here are interpretation laid on top of that result.
A second qualification runs the other way. Human stroboscopic assessment has its own known variability, and interrater reliability differs by which feature is being rated, a limitation documented well before generative models entered the picture ([Mehta, Current role of stroboscopy in laryngeal imaging, 2012]). The point is not that clinician judgment is unassailable. It is that this study tested a task clinicians do not actually perform.
What Is Documented and What Is Not
The concerns raised around synthetic medical imaging, including fraudulent claims, diagnostic deception, and manipulated figures in the literature, are concerns researchers have articulated, not documented harms with incidence data behind them. That distinction matters and should not be blurred.
Research image integrity is the one area with hard numbers, and they predate this technology entirely. A screening of 20,621 papers across 40 journals published between 1995 and 2014 found problematic figures in 3.8% of them, with at least half showing features suggestive of deliberate manipulation ([Bik, The Prevalence of Inappropriate Image Duplication in Biomedical Research Publications, 2016]). That analysis examined image duplication, not AI generation. Generative tools did not create this problem, but they arrive on top of a baseline that was already substantial.
On the other side of the ledger, the case for synthetic images as training data remains a mechanistic expectation rather than a demonstrated result. This study tested whether images fool human observers, not whether they improve any downstream model.
Key Takeaways
In a 2026 study, 114 clinicians identified authentic laryngeal images 70.9% of the time but synthetic ones only 57.8% of the time (p<0.001).
Perceptual realism plateaued early: accuracy fell from 79.6% to 59.1% between the first two training checkpoints and improved little afterward, bottoming at 44.8%.
Participants reading on a computer outperformed those on a phone (63.9% vs 56.2%, p=0.014), though device was self-selected rather than randomized.
The study measured perceptual realism, not diagnostic accuracy, and its published level of evidence is listed as not applicable.
Because the survey used still frames while stroboscopy is a modality for assessing vibration, neither the forgery risk nor the training-data benefit extends automatically to full clinical examinations.
FAQ
Can doctors tell AI-generated medical images from real ones?
Not reliably, based on current evidence. In the laryngeal study, overall accuracy was 59.9%, and synthetic images were correctly flagged only 57.8% of the time. A separate radiology study reported accuracies of 70% and 75% on two datasets once readers knew synthetic images were present, but those figures come from a different design and are not directly comparable to the laryngeal numbers. Neither study describes dependable detection.
Is this the same as the deepfake X-ray study?
No. They are independent studies using different generative models, different reader groups, and different task designs, published five months apart. Their accuracy figures are not directly comparable. What they share is the question being asked.
Does this mean AI improves laryngoscopy diagnosis?
No. The study did not examine diagnostic performance at all. It measured perceptual realism, not downstream model accuracy. An image that fools a human observer has not been shown to make a classifier better, and whether synthetic laryngeal images help train diagnostic systems remains an open question.
References
- Morrison DA, Mohanty AS, Sulica L, Khosravi P, Rameau A. Human Evaluation of Synthetic Videostroboscopic Laryngeal Images Generated Using StyleGAN3. Laryngoscope. 2026.
- Tordjman M, Yuce M, Ammar A, et al. The Rise of Deepfake Medical Imaging: Radiologists’ Diagnostic Accuracy in Detecting ChatGPT-generated Radiographs. Radiology. 2026;318(3):e252094.
- Mehta DD, Hillman RE. Current role of stroboscopy in laryngeal imaging. Curr Opin Otolaryngol Head Neck Surg. 2012;20(6):429-36.
- Bik EM, Casadevall A, Fang FC. The Prevalence of Inappropriate Image Duplication in Biomedical Research Publications. mBio. 2016;7(3).
Joonpyo Hong, MD is a board-certified otolaryngologist practicing in Korea. This article reflects his clinical interpretation of published research and does not constitute individual medical advice.
For more articles:
https://curiousmd.com/how-ai-voice-cloning-works/
https://curiousmd.com/ai-laryngeal-cancer-detection/
https://curiousmd.com/neuralink-voice-speech-first/
Link out to:
https://www.rsna.org/news/2026/march/chatgpt-generated-radiographs — RSNA’s own write-up of the deepfake radiograph study, including the lead author’s account of the visual tells that give synthetic X-rays away.
https://nvlabs.github.io/stylegan3/ — NVIDIA’s project page for StyleGAN3, the model used to generate the laryngeal images. The comparison videos show what the architecture changed.
https://www.ncbi.nlm.nih.gov/books/NBK567774/ — A plain-language StatPearls overview of videostroboscopy: how the technique samples vocal fold vibration, and where its limits lie.
