A radiology worklist notification indicates that a possible pulmonary embolism has been detected and that the study should be prioritized for review. On opening the examination, the radiologist finds that the AI triage tool has generated an alert but has provided no probability estimate, no confidence metric, and no indication of which images or anatomical regions produced the classification. The system outputs a binary decision only, without quantification of the certainty underlying that determination. This describes the operational behavior of most FDA-cleared AI devices currently deployed in clinical radiology rather than that of a research prototype. These systems function as triage tools that flag examinations for priority review or route them to specialized worklists, and most provide nothing beyond a binary classification, since that is the output current regulatory clearances address. When such a system produces an incorrect classification, which is expected given that no algorithm achieves perfect accuracy, the clinician has no means of assessing the confidence associated with that prediction. A normal anatomical variant such as a growth plate or nutrient canal misclassified as a fracture generates the same binary notification as an unambiguous displaced femoral fracture with cortical disruption, notwithstanding the substantial difference in diagnostic certainty that should attach to the two classifications.

The Binary Triage Problem

The predominant category of FDA-cleared AI devices deployed in clinical radiology comprises triage tools that analyze examinations in the background and generate binary classification alerts, such as detection of intracranial hemorrhage, large vessel occlusion, pneumothorax, or rib fracture. These systems reprioritize flagged examinations on the radiologist worklist or route them to specialized reading queues. They provide no quantitative indication of the confidence underlying the classification. Radiologists reviewing flagged cases receive no probability estimate, no uncertainty quantification, and frequently no visualization indicating which images, regions, or imaging features produced the alert. Internally, the model has computed a continuous-valued score that was compared against a fixed threshold to generate the binary output, but neither the threshold nor the score is exposed to the user. Consequently, a case that exceeded the threshold marginally, for example with an internal score of 0.51 on a 0 to 1 scale, receives the same positive classification as a case with an internal score of 0.99, despite these values representing substantially different degrees of model certainty.

The clinical consequence is that every positive alert is treated as equally urgent irrespective of the model's confidence. Clinicians cannot distinguish reliable alerts warranting immediate attention from marginal classifications that exceeded the threshold by a small margin. When the system incorrectly identifies a normal anatomical structure as a fracture, that false positive generates the same notification and workflow interruption as a true positive requiring emergent orthopedic consultation. Absence of confidence information therefore prevents prioritization on the basis of classification certainty.

Clinical Implications

The implications of absent uncertainty quantification may be illustrated with two representative cases encountered by an automated pneumothorax detection system. In Case A, a large tension pneumothorax is present on an upright posteroanterior chest radiograph, with mediastinal shift, lung collapse, and tracheal deviation. The model's internal classification score is 0.95, reflecting high certainty on the basis of unambiguous imaging features. In Case B, a skin fold artifact is present on a supine portable chest radiograph obtained in an intensive care unit patient, in which the linear density resembles the pleural line displacement characteristic of pneumothorax. The model's internal score is 0.52, marginally above the classification threshold of 0.50. The output presented to the radiologist is identical in both cases: pneumothorax detected, examination moved to the top of the list. Without access to information distinguishing these scenarios, clinicians must treat both alerts equivalently. This may produce inappropriate workflow interruption, for example suspension of a procedure to review Case B, which represents artifact, while the absence of prioritization information may delay review of Case A, which requires immediate intervention. Moreover, accumulation of false positive alerts associated with low-confidence classifications contributes to alert fatigue, which may reduce responsiveness to all alerts from the system, including high-confidence detections of clinically critical findings.

Definition of Uncertainty Quantification

Uncertainty quantification (UQ) in medical AI is the computational analog of a radiologist communicating diagnostic doubt, that is, the capacity to convey not only a classification but also a quantitative assessment of the reliability of that classification. Radiologists routinely communicate varying degrees of certainty through report language, ranging from definitive statements to qualified interpretations such as suspicious for or compatible with, to explicit acknowledgment of uncertainty through a stated differential or a recommendation for additional imaging. Uncertainty quantification permits AI systems to communicate confidence in an analogous manner. Rather than a binary flag indicating presence or absence of a finding, a UQ-enabled triage system can provide a quantitative confidence metric, a probability distribution over candidate diagnoses, an estimate of classification reliability, and a visual overlay indicating the image regions contributing most to the prediction. Uncertainty in this context represents the reliability of the model's prediction: high uncertainty indicates that the model has insufficient information to generate a reliable classification and that the prediction should be treated with caution, whereas low uncertainty indicates that the prediction is more likely to be accurate.1

The principal operational benefit is that a triage system can distinguish between a flagged case about which it is uncertain, which warrants careful review, and a flagged case about which it is confident, which warrants immediate action.

Aleatoric and Epistemic Uncertainty

Identifying the source of uncertainty determines whether it can be reduced.

  • Aleatoric uncertainty (data uncertainty): irreducible noise inherent in the data. Some chest radiographs are visually indistinguishable but correspond to different diagnoses given differing clinical history. Additional training data will not eliminate this component, since the ambiguity is a property of the input.
  • Epistemic uncertainty (model uncertainty): uncertainty attributable to gaps in the model's training distribution, that is, insufficient exposure to comparable examples. This component can be reduced by training on more varied data.1

High epistemic uncertainty indicates that the model has encountered a case unlike those represented in its training data. This is actionable information, since it identifies the cases in which human review is most likely to be necessary.

Reliability as a Distinct Objective from Accuracy

Limitations of Accuracy as a Sole Criterion

For approximately a decade, development in machine learning has been organized around improvements in benchmark accuracy, whether measured by ImageNet performance or by scores on aggregate reasoning benchmarks. Predictive performance has served as the primary criterion of progress.

That emphasis is shifting, and we regard the shift as a precondition for adoption in healthcare. Deployment of large language models and autonomous vision systems in high-stakes clinical settings has demonstrated a specific limitation. These models are frequently accurate and fluent yet poorly calibrated: they produce confident output when incorrect, do not reliably identify inputs that lie outside their training distribution, and do not communicate the absence of sufficient information to answer. Uncertainty quantification, including distribution-free methods such as conformal prediction, addresses this limitation directly.

Conformal Prediction and Distribution-Free Guarantees

Beyond heuristic calibration methods such as temperature scaling, conformal prediction (CP) has received increasing attention. Unlike Bayesian approaches, which depend on specified priors, CP provides finite-sample coverage guarantees. Rather than a single label, CP returns a prediction set that contains the true label at a specified rate.

However, standard CP depends on the exchangeability assumption, namely that test data are statistically comparable to the calibration data. In deployment, this assumption is frequently violated.

Conformal Methods Under Distribution Shift

Recent variants extend CP to the distribution shift setting. Weighted conformal approaches, for example, adjust the prediction set according to the degree to which an input departs from the calibration distribution.

This is relevant in medicine, in which the deployed population rarely corresponds exactly to the calibration set. A model calibrated at one institution may encounter a different case mix, scanner, or referral pattern at another. Weighted conformal methods preserve coverage under such shift by assigning greater weight to calibration cases resembling the new input.

Why Calibration Is Not Sufficient

It may be argued that calibration addresses this problem, in that a model reporting 80% should be correct in 80% of such cases. Calibration is valuable but is not sufficient, for the following reasons.

Calibration establishes that predicted probabilities correspond to observed frequencies in the calibration dataset. If a model predicts a 70% probability of pneumonia for 100 cases and 70 of those cases have pneumonia, the model is calibrated.

However, calibration is a population-level property. It does not establish whether a particular prediction is reliable. A perfectly calibrated model may be either highly uncertain or highly confident about individual cases in a manner that the aggregate probability does not reveal.4

Consider two calibrated pneumonia models:

  • Model A: trained on numerous cases of Klebsiella pneumonia. Presented with a Klebsiella case, it predicts a 70% probability of pneumonia with low uncertainty (0.05).
  • Model B: not trained on Klebsiella pneumonia. Presented with the same case, it predicts a 70% probability of pneumonia with high uncertainty (0.6).

Both models are calibrated and both output 70%. Model B is nevertheless substantially less reliable for this case, because the presentation is not represented in its training distribution. Only uncertainty quantification distinguishes the two.

Dependence of Calibration on Prevalence

Calibration is additionally not portable across settings. A model calibrated on a population with 20% disease prevalence requires recalibration for deployment in a setting with 60% prevalence.

Vendor claims regarding calibration should therefore be validated on the local population, and even following such validation, uncertainty quantification remains necessary to determine which individual predictions are reliable.

Methods for Uncertainty Quantification

Several approaches to quantifying uncertainty in deep learning models are available. The most commonly used are described below.

1. Ensemble Methods

Multiple models are trained on the same task with different initializations or architectures, and each input is evaluated by all models. Agreement among the models, for example predictions in the range of 80–85%, indicates low uncertainty. Disagreement, for example predictions ranging from 30% to 90%, indicates high uncertainty.

Advantages: conceptually straightforward and empirically effective.
Disadvantages: computationally expensive, since multiple models must be trained and executed.

2. Bayesian Approximation by Monte Carlo Dropout

The same model is evaluated multiple times on the same input with different neurons randomly disabled on each pass. This approximates an ensemble without training multiple models, and the variance across passes serves as an uncertainty estimate.

Advantages: requires a single trained model and is comparatively efficient.
Disadvantages: still requires multiple forward passes at inference.1

3. Evidential Deep Learning

The model is trained to accumulate evidence for each class from the image features, with greater accumulated evidence corresponding to lower uncertainty. This approach produces an uncertainty estimate in a single forward pass.

Advantages: efficient at inference and theoretically principled.
Disadvantages: more complex to implement and less extensively evaluated in medical imaging.1

Applications of Uncertainty Quantification

Selective Referral for Expert Review

The most direct application is threshold-based routing: cases for which uncertainty exceeds a specified threshold are directed to a radiologist for careful review, whereas high-confidence cases may be assigned lower priority or subjected to abbreviated review.

This produces a hybrid workflow in which the model handles routine cases and human expertise is applied selectively to cases that the model identifies as outside its reliable operating range.

Improvement in Model Performance

Uncertainty quantification may also improve diagnostic performance directly. In segmentation tasks, identification of the voxels about which the model is uncertain permits targeted refinement.

In one study, incorporation of uncertainty into brain tumor segmentation improved Dice coefficients by 3.15% for enhancing tumor and 0.58% for necrotic tumor, differences that may be clinically relevant for treatment planning.3

Active Learning

Model retraining as additional data become available is constrained by the cost of expert annotation, which raises the question of which cases should be annotated.

Cases associated with the highest model uncertainty are the most informative for subsequent training. Uncertainty quantification therefore permits targeted data collection rather than random sampling of cases for annotation.1

Detection of Distribution Shift and Bias

High epistemic uncertainty may indicate that a model is being applied outside its training distribution, which is frequently associated with disparate performance.

For example, a pneumonia model trained on pediatric examinations and deployed on adult patients would be expected to exhibit high epistemic uncertainty on adult cases, indicating that the model is unreliable in that population. In the absence of uncertainty quantification, this would be identified only through observed clinical errors.

As discussed in our article on algorithmic bias, models trained on non-representative populations propagate that composition to downstream predictions. Uncertainty quantification may identify cases in which a model is uncertain specifically because the patient belongs to a group underrepresented in training.

Questions for Vendors

The following questions are recommended when evaluating AI triage tools for clinical deployment.

1. Does the tool provide confidence or uncertainty information with each alert?

Most currently cleared devices do not, providing binary output only.

Minimum acceptable: categorical confidence levels displayed with each alert.
Preferable: quantitative uncertainty scores that can be used to define routing rules.
Optimal: uncertainty estimates accompanied by indication of the image features contributing to the alert.

Inadequate response: a statement that the model has 95% sensitivity and that all alerts should be treated as equally urgent. Sensitivity is a population-level property and does not establish which individual alerts are reliable.

2. By what method is uncertainty quantified?

Where uncertainty is provided, the method should be specified, whether ensemble, Bayesian approximation, evidential, or conformal. It should be noted that vendors frequently equate the output of the final softmax layer with a probability, which it is not, or describe a calibrated probability as the model's confidence, which it also is not. These distinctions are material to interpretation.

3. What is the positive predictive value stratified by confidence level?

Aggregate PPV is insufficient. The PPV among high-confidence alerts and among low-confidence alerts should be requested separately. Inability to provide stratified values indicates that the uncertainty estimates have not been validated.

Example of an adequate response: high-confidence alerts had a PPV of 65% and low-confidence alerts a PPV of 12% in the validation study.

4. Can alert behavior be configured by confidence level?

It should be possible to specify that high-confidence alerts generate an interruptive notification while low-confidence alerts are added to the worklist without interruption. Vendors should provide the configuration mechanisms required to adapt this behavior to local practice.

5. How does the tool handle out-of-distribution inputs?

Specifically, whether uncertainty increases when a model trained on adults encounters a pediatric examination, and whether uncertainty increases in the presence of artifact or atypical positioning.

A statement that the model performs equally well across all cases is not plausible and indicates that out-of-distribution behavior has not been characterized.

Uncertainty Quantification and Clinical Trust

Clinical AI is subject to a dual requirement: models must be accurate and must also represent the limits of their competence. A model that produces confident incorrect output is more hazardous than one that produces no output, because it presents no signal that review is required.

Uncertainty quantification addresses this requirement by providing a mechanism for the system to communicate doubt. An expressed uncertainty is not an indication of failure; it identifies cases in which the input lies outside the region in which the model performs reliably and human review is therefore indicated.

This form of communication is what permits sustainable integration into clinical workflows. Radiologists do not require that a tool be infallible; they require that its reliability be legible on a case-by-case basis.

Implications for Deployment

As discussed in our article on deployment failure in medical AI, most medical AI does not achieve sustained clinical use, and insufficient trust is a principal contributing factor.

Models that do not represent uncertainty fail in deployment for identifiable reasons:

  • They generate false positives that are indistinguishable from true positives
  • Radiologists cannot separate reliable from unreliable predictions
  • No mechanism exists for routing uncertain cases differentially
  • Each confident error reduces subsequent reliance on the system

Models incorporating uncertainty quantification, by contrast, can:

  • Identify inputs outside their reliable operating range
  • Route cases according to estimated confidence
  • Preserve calibrated clinician trust by representing their own limitations
  • Direct expert review to the cases in which it is most informative

Conclusion

Most FDA-cleared AI triage tools currently provide binary alerts without confidence scores, uncertainty estimates, or, in many cases, any display of the finding that produced the alert. Every alert is presented as equally urgent, irrespective of whether the model's internal score exceeded the classification threshold by a wide or a marginal margin.

This design is limiting. Where confident predictions cannot be distinguished from marginal ones, alert fatigue is the expected outcome, and clinicians respond by disregarding alerts or by applying uniform skepticism to all of them, which reduces the value of the tool.

Uncertainty quantification addresses this limitation. The methods are technically established and have been validated in research settings. What remains is adoption, specifically requiring that vendors provide confidence information with triage alerts and declining to accept systems that present every alert as equally reliable.

A triage tool that distinguishes a high-confidence from a low-confidence detection of intracranial hemorrhage permits differential workflow routing. High-confidence alerts warrant immediate attention, whereas low-confidence alerts may be reviewed without interruption of ongoing work. Calibrated trust is preserved because the system represents its own limitations.

The principal hazard is therefore not a model that is occasionally incorrect, but a model that is incorrect with the same apparent confidence it exhibits when correct. A triage tool that can identify the cases about which it is uncertain is more useful, and safer, than one that cannot.


Key Takeaways

  • Most triage tools output a binary classification: FDA-cleared AI devices typically provide a detected or not detected flag without a probability estimate, a confidence score, or, in many cases, any display of the finding.
  • Binary alerts do not distinguish marginal from unambiguous cases: a detection with an internal score of 0.51 receives the same notification as one with an internal score of 0.99, and the clinician cannot distinguish between them.
  • Alert fatigue is a predictable consequence: where a high proportion of flags are false positives and all are presented identically, clinicians deprioritize or disregard alerts, which reduces the value of the tool.
  • Uncertainty quantification permits differential routing: categorical or quantitative confidence levels allow high-confidence alerts to be prioritized while low-confidence alerts are reviewed without interrupting ongoing work.
  • Two components of uncertainty: aleatoric uncertainty reflects irreducible ambiguity in the data, whereas epistemic uncertainty reflects gaps in the training distribution and can be reduced with more diverse training data.
  • Calibrated trust requires representation of uncertainty: a tool that distinguishes high-confidence from low-confidence detections is more useful, and safer, than one that presents every alert identically.
  • Performance data to request from vendors: whether confidence is reported, whether the detected finding is displayed, the positive predictive value stratified by confidence level, and whether alert behavior can be configured by confidence.

References & Further Reading

  1. Faghani S, Moassefi M, Rouzrokh P, Khosravi B, Baffour FI, Ringler MD, Erickson BJ. Quantifying Uncertainty in Deep Learning of Radiologic Images. Radiology. 2023;308(2):e222217. doi:10.1148/radiol.222217
  2. Dohopolski M, Chen L, Sher D, Wang J. Predicting Lymph Node Metastasis in Patients with Oropharyngeal Cancer by Using a Convolutional Neural Network with Associated Epistemic and Aleatoric Uncertainty. Phys Med Biol. 2020;65(22):225002. doi:10.1088/1361-6560/abc2d0
  3. Lee J, Shin D, Oh SH, Kim H. Method to Minimize the Errors of AI: Quantifying and Exploiting Uncertainty of Deep Learning in Brain Tumor Segmentation. Sensors (Basel). 2022;22(6):2406. doi:10.3390/s22062406
  4. Guo C, Pleiss G, Sun Y, Weinberger KQ. On Calibration of Modern Neural Networks. Proceedings of the 34th International Conference on Machine Learning. 2017;1321-1330.
  5. Abdar M, Pourpanah F, Hussain S, et al. A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges. Inf Fusion. 2021;76:243-297. doi:10.1016/j.inffus.2021.05.008
  6. Ozdemir O, Russell RL, Berlin AA. A 3D Probabilistic Deep Learning System for Detection and Diagnosis of Lung Cancer Using Low-Dose CT Scans. IEEE Trans Med Imaging. 2020;39(5):1419-1429. doi:10.1109/TMI.2019.2947595
  7. Rajaraman S, Ganesan P, Antani S. Deep Learning Model Calibration for Improving Performance in Class-Imbalanced Medical Image Classification Tasks. PLoS One. 2022;17(1):e0262838. doi:10.1371/journal.pone.0262838