Artificial intelligence (AI) tools marketed for radiologist efficiency are now widely available, and reported time savings are frequently cited in vendor materials and trade publications. However, the magnitude of the reported effect varies considerably across sources, and the underlying studies differ substantially in design, setting, and outcome measure. In this article, we review the published evidence for AI-associated efficiency gains in radiology and characterize what has been demonstrated, under what conditions, and with what degree of confidence.
The principal conclusion is that the evidence is genuine but considerably more conditional than promotional summaries indicate. The proposition that AI improves radiology efficiency is not a single claim. It comprises at least four distinct claims, supported by four different classes of evidence, with correspondingly different confidence levels. Conflating them is a common source of unmet expectation following procurement.
Taxonomy of AI Efficiency Tools
The term "AI efficiency tool" is applied inconsistently, and this imprecision permits strong results in one product category to be cited in support of weaker claims in another. We therefore distinguish four categories that are currently marketed and studied.
Report drafting assistants accept an image, typically a radiograph or CT examination, and generate a draft report that the radiologist accepts, edits, or rejects. In each case the radiologist edits generated text rather than composing from an empty template, and the efficiency gain is expected to scale with the proximity of the draft to the final report.
Triage and worklist tools do not modify the report. They reorder the reading queue so that examinations with a higher predicted probability of containing a finding of interest, such as stroke, pneumothorax, or pulmonary embolism, are presented before examinations ordered by arrival time.
Protocol selection and quantification tools perform narrow, well-specified tasks such as MRI protocol assignment or nodule measurement. These tasks involve no free-text generation and no open-ended reasoning.
General-purpose vision-language models (VLMs) perform open-ended diagnostic reasoning across arbitrary pathology. This category attracts the greatest attention and, on the available evidence, warrants the greatest caution.
Notably, these categories do not exhibit comparable reliability, and the difference is large. Studies of narrow, closed-set tasks such as protocol selection and single-finding detection report accuracy in the range of 84–98%. By comparison, an evaluation of a general-purpose vision-language model performing open-ended diagnosis across 230 emergency department cases reported 35% accuracy for pathology identification and an overall hallucination rate of 47%, indicating that the model described findings that were not present in nearly half of cases. It is worth noting that the same model performed well on the more constrained components of the task, achieving 100% accuracy for imaging modality identification and 87% for anatomical region.
This pattern is not internally inconsistent. It reflects the difference between constrained and open-ended tasks, and it implies that a reported accuracy figure is largely uninterpretable without knowledge of the task category to which it applies. A worklist prioritization tool and a free-text drafting VLM do not make equivalent claims and should not be evaluated against a common evidentiary standard.
Efficiency Evidence by Category
Report drafting: modest effects, possibly concentrated in complex cases
The most rigorously designed study identified used a within-reader crossover design in which three radiologists reported 50 chest radiographs twice, once unassisted and once with AI drafting assistance, with direct measurement of the difference. Overall reporting time decreased by 7.8%, which did not reach statistical significance at this sample size. Notably, in the subgroup of complex reports, the reduction was 18.3% and did reach statistical significance, with no measured decrement in report quality.
This finding is potentially important, as it suggests that the benefit of drafting assistance may concentrate in more difficult cases rather than in routine ones. However, it derives from a single small study, and subgroup effects of this kind are frequently attenuated on replication. Additional studies are directionally concordant at differing magnitudes. One single-site study reported a decrease in reporting time from approximately 9 minutes to 6.75 minutes per case, corresponding to a reduction of approximately 24%. A multicenter study of AI-generated brain MRI impressions reported a decrease in reading time from 61 to 53 seconds, with a larger effect among junior radiologists than among senior radiologists. If this interaction is confirmed, it may have implications for deployment in training programs and in departments operating below target staffing.
One point warrants explicit qualification. A frequently cited 2025 study in Nature Medicine of a vision-language model designated Flamingo-CXR, conducted across 27 radiologists and 606 cases, reported that AI-assisted reports were rated equivalent or superior to AI-only reports in most comparisons. This result is often presented as efficiency evidence. It is not. The authors state explicitly that reporting time was not measured and that doing so "warrants another carefully designed human study." The absence of a time endpoint in a study of this scale is relevant context when evaluating smaller studies that report definitive time savings.
Summary for drafting tools: the available evidence supports a modest improvement in reporting time, on the order of single digits to low double digits in percentage terms, potentially concentrated in more complex or time-consuming cases. Specific figures quoted by vendors should be treated as provisional until reproduced in the local workflow. It should also be noted that most of this evidence concerns chest radiography and may not generalize to cross-sectional modalities such as CT.
Triage and worklist prioritization: the strongest evidence, with an important qualification
The largest reported effects, and the most consequential caveat, are found in this category. A 2020 simulation study modeling AI-based worklist reordering against a first-in-first-out baseline reported turnaround-time improvements of 32–56%, with larger gains for pneumothorax and smaller gains for pleural effusion. A 2024 study of AI-assisted detection of aortic dissection reported a 68% reduction in scan-to-assessment time for confirmed positive cases.
Both are simulation studies rather than prospective clinical deployments. This is a defensible design choice, since prospective evaluation of an unvalidated triage algorithm carries patient risk. However, it means that the reported values estimate performance in a simulated replay of historical data rather than in an operating department, and the weight assigned to them should be adjusted accordingly.
This literature also illustrates how imprecise figures propagate. A claim of "77% turnaround-time reduction, 99% specificity" for AI worklist triage circulates in secondary sources and vendor materials. Tracing this claim to its apparent origin identifies the 2020 simulation described above, which reports 32–56% improvement and contains no specificity figure of 99%. The 77% and 99% values are not supported by the primary source. Requests for primary sources rather than secondary summaries are therefore appropriate when such figures inform procurement.
A further consideration applies to all prioritization tools. Where turnaround time improves for examinations containing the index finding, turnaround time for all remaining examinations is presumably prolonged, and those examinations are not necessarily normal. They may contain renal laceration or hepatic metastases, findings of comparable clinical importance. Moreover, the timely establishment of a negative result carries operational value in an emergency department constrained by room availability. Reported gains restricted to the index finding therefore do not describe the net effect on the queue.
The strongest real-world evidence in this domain derives from a 2026 study of more than 1.1 million examinations at a tertiary hospital in China, comparing outcomes before and after implementation of a scheduling and dispatch platform. Examination volume increased by 66% without an increase in staffing, and the platform reduced median time from order to appointment by 34–72% and from order to examination by 35–51%, with all differences reaching statistical significance. However, the outcome measured is the rate at which a patient progresses through scheduling and dispatch, not the rate at which a radiologist interprets and reports an examination once it appears on the worklist. These are related but distinct questions, and promotional material frequently does not distinguish between them.
Summary for triage tools: simulated evidence for reduced time to assessment of urgent findings is consistent and comparatively strong. Real-world evidence for improved patient throughput and scheduling is also strong. What remains absent is real-world evidence that triage tools measurably accelerate the radiologist's interpretation of the examination, as distinct from delivering the examination to the radiologist earlier.
Workforce capacity and burnout: directionally encouraging, sparsely evidenced
A 2025 review synthesizing 22 studies of AI in the context of the radiologist shortage reports an approximate 53% reduction in AI-assisted workload. However, the same review reports that per-image interpretation time in the underlying studies increased, from 3–4 seconds to 6–7 seconds. These figures are not necessarily contradictory, as the former plausibly reflects aggregate case throughput and the latter the overhead of reviewing an AI-generated second read. Nevertheless, the review does not provide sufficient detail to reconcile them with confidence, and department-specific data should be requested before either figure is relied upon.
The evidence regarding burnout is thinner. One review cites a pilot deployment of an AI scribe associated with a 40% reduction in self-reported burnout. This derives from a single uncorroborated pilot study and should be regarded as hypothesis-generating rather than as a basis for programmatic decisions.
Summary for workforce and burnout claims: directionally plausible but substantially under-evidenced. Where a vendor cites a specific burnout reduction percentage, the number of contributing sites and radiologists should be requested; in the current published literature, that number is frequently one.
An Unmeasured Quantity: The Cost of Error Correction
The most consequential gap in this evidence base concerns what is not measured. Every efficiency study identified measures time saved. None measures the time required to detect and correct the model's errors.
General-purpose diagnostic models are known to hallucinate at a high rate on open-ended tasks; the 47% rate cited above is consistent with results reported whenever such models are applied beyond narrow, well-defined tasks. If a radiologist saves two minutes during drafting but expends three minutes verifying a hallucinated finding, the net effect is negative. At present, no published study measures both quantities within the same workflow. Efficiency research and diagnostic accuracy research currently constitute largely separate literatures. Until a study measures both jointly, every time-savings estimate in this article, and every such estimate in a vendor's materials, should be interpreted as time saved conditional on the absence of an error requiring correction. This is a materially weaker claim than it is generally presented to be.
Regulatory Clearance Is Not Evidence of Efficiency
The number of FDA-authorized AI medical devices has grown rapidly, and radiology accounts for the majority of authorizations: more than 1,450 authorizations in total through the end of 2025, of which radiology represents approximately three-quarters. This indicates an active regulatory pathway and sustained vendor engagement with it. It does not constitute evidence of efficiency benefit. A 510(k) clearance establishes substantial equivalence to an existing predicate device, which is a safety and equivalence determination rather than a demonstration of improved departmental throughput. Citation of FDA clearance in support of a time-savings claim is a category error.
An unresolved tension exists between the pace of regulatory clearance and a separate body of literature describing clinical validation of AI as lagging behind technical capability. These two observations have not been fully reconciled, and doing so would require regulatory-science expertise that is not typically available within a clinical department. The appropriate response is to acknowledge the tension rather than to treat clearance as though it settled the question.
Implications for Radiologists
None of the foregoing argues against the use of these tools. The directional signal toward genuine time savings, particularly for drafting assistance on complex cases and for triage prioritization of urgent findings, appears across independent studies from independent groups, and such convergence is unlikely to be coincidental. Several considerations nevertheless apply before a workflow is modified on the basis of a vendor efficiency claim.
First, the product category should be identified and the tool held to the evidentiary standard that category has established. A narrow protocol automation tool achieving greater than 90% accuracy on a well-defined task represents a substantially different proposition from a general-purpose model reasoning over open-ended pathology. A strong result in one category does not support a weak claim in another.
Second, specific figures should be traced to their primary sources. As demonstrated above, a widely repeated percentage may have become detached from a source that does not support it. Where a specific value materially affects a decision, the underlying study rather than a secondary summary should be consulted.
Third, claims that imply the elimination rather than the reduction of verification work warrant particular scrutiny. Tools that draft or propose report content still require the radiologist's judgment, and studies reporting time savings generally measured expert radiologists performing that verification carefully. Omission of verification because the model is usually correct constitutes the automation bias that the safety and medicolegal literature identifies as a principal risk, precisely because usually correct and correct in the present case are not equivalent.
Finally, we have argued elsewhere on this site that model confidence, and conversely uncertainty, is an informative output that is rarely used. If high-confidence outputs can be identified reliably and are correct at a high rate, the benefit of each of the tool categories described above may be amplified by suppressing low-confidence outputs that would otherwise consume verification effort.
Implications for Department Management
Local piloting with locally measured endpoints is preferable to reliance on published values. The most consistent finding across this literature is the degree to which results depend on setting; one review reported both substantial efficiency gains and the absence of any turnaround-time improvement for comparable tools within the same document. Case mix, existing workflow, staffing model, and patient population are likely to determine the local result more than any published study will. A well-conducted 60-day pilot against the department's own baseline is more informative than a vendor white paper.
The procurement question should be separated from the problem-definition question. A tool that accelerates patient scheduling and dispatch addresses a different problem from one that accelerates radiologist interpretation, and a tool that reduces reported burnout addresses a third. The bottleneck should be characterized before candidate tools are evaluated against it.
Cost-effectiveness should be assessed explicitly, since the efficiency literature does not address it. None of the studies reviewed weighs time saved against licensing, integration, and maintenance costs, as these lie outside the measured outcomes. A tool that saves five minutes per case is not necessarily worth adopting if its licensing cost exceeds the value of that time under the local staffing model. This calculation is distinct from any efficiency percentage and must be performed separately.
Regulatory clearance should not be treated as an answer to the efficiency question, as it is not designed to be one. Vendors should be asked directly for efficiency evidence and for the setting in which it was obtained, specifically whether the study was simulated or prospective and whether it was single-site or multi-site. Clearance and demonstrated time savings are separate claims.
Directions for Future Work
The most useful contribution to this evidence base over the next several years would be a prospective, multi-site, real-world trial measuring radiologist interpretation time and AI error-correction burden concurrently, on the same cases. In the absence of such a study, every efficiency claim in this literature, including those reviewed here, is directionally informative and individually insufficient to justify a departmental workflow change without local pilot data.
This is not an argument for deferring adoption. It is an argument for conducting a small, well-measured local pilot before scaling, for asking vendors more specific questions than the aggregate time saved, and for recognizing that accountability for the content of the final report remains with the interpreting radiologist rather than with the assisting tool.
This article draws on a scoping review of 17 peer-reviewed and preprint sources examining AI and vision-language model tools for radiologist workflow efficiency, published between 2020 and 2026. The full review, including evidence tables, source verification tiers, and methodology, is available on request.