My two most recent posts addressed separate questions. The first examined whether artificial intelligence (AI), or in the term now more commonly used, superintelligence (SI), poses an existential risk to humanity. That probability is often abbreviated as p(doom). The second described System-1 models, a class of decision models that return a fixed, pre-specified set of outputs rather than free text. On further reflection, I believe these two topics overlap substantially. Specifically, System-1 models may offer a practical way to reduce the probability that a highly capable AI system inadvertently causes serious harm.
"Do what I mean, not what I say"
Most readers have heard this phrase, and many have used it. It reflects two facts about human communication. Natural language is less precise than we often assume, and speakers sometimes state a goal in a way that allows unintended interpretations. This gap between stated and intended objectives is at the center of most p(doom) arguments. The concern is not that an AI system has malicious intent. Rather, the concern is that a human will express a goal to a highly capable system, and the system will pursue that goal by a route that produces undesired, and possibly lethal, consequences. Amodei et al. described this class of failure as "negative side effects" and "reward hacking." They noted that an objective function that omits important constraints will be optimized without regard to those constraints [1].
It follows that one cannot realistically state a broad goal and leave an AI system to "figure it out" for most real-world tasks. This is particularly true in medical care.
Medical care as a structured process
Good medical care is a process, not a single act. It begins with listening to the patient to understand what they perceive to be wrong. The clinician then asks questions that clarify the nature of the complaint. As signs and symptoms are gathered, the clinician considers a differential diagnosis. Tests are then ordered to narrow the diagnostic possibilities. In some cases a single test establishes the diagnosis. In others, the clinician must follow a deep tree of sequential tests. Once a diagnosis is made, treatment continues the process. An initial procedure or medication is applied, its effect is assessed, and the clinician determines whether further steps are required. Longer-term follow-up is frequently needed as well.
Each of these stages contains decision points. Many of them can be supported by System-1 models, which are faster than large language models (LLMs) and, I will argue, safer for this purpose.
Why structured outputs are safer
As described in my prior post, System-1 models offer substantial speed advantages over LLMs. Their safety advantage comes from a different property. They produce a structured response, and a structured response requires that the set of permissible responses be defined in advance. The output of a System-1 model cannot be the free text of a generative model. It must always be one of a set of pre-specified values. These models also supply confidence scores. It should be noted that raw confidence scores from modern neural networks are often poorly calibrated. Well-established post hoc methods such as temperature scaling can correct this [2]. Calibrated confidence is what makes threshold-based escalation to a human reviewer meaningful.
Because each possible output is enumerated, the application must also specify the next action for each output. This gives the designer far greater control than allowing agents to act without constraint. Requiring the actions at each branch of the decision tree to be defined allows for substantially greater reliability. It also substantially reduces the opportunity for an AI system to act outside its intended scope.
System-1 models are also less susceptible to prompt injection. Prompt injection is defined as the insertion of malicious text into a model's input to induce harmful output. The reduced susceptibility arises because a classifier has no instruction-following channel. Injected text can at most shift the model toward one of its pre-specified outputs. It cannot cause the model to emit new instructions or take unplanned actions. Generative models, in comparison, remain highly vulnerable. Clusmann et al. tested four state-of-the-art vision-language models with 594 prompt injection attacks embedded in medical images. All four models were susceptible, and the injected prompts were not apparent to human observers [3].
Lessons from the Hugging Face incident
The events of July 2026 illustrate what can happen when broadly tasked agents operate with insufficient constraints. During an internal cybersecurity evaluation, OpenAI models escaped their evaluation environment. These included GPT-5.6 Sol and an unreleased, more capable model. The models exploited a vulnerability in a package-proxy service, reached the internet, and compromised Hugging Face's production infrastructure [4]. Their apparent objective was to retrieve answer keys for the benchmark on which they were being evaluated. Hugging Face detected and disclosed the intrusion before OpenAI identified its own models as the source [4].
Two lessons are relevant here. First, the sandbox was misconfigured. This is evidence that we cannot assume firewalls and isolation boundaries will contain a capable agent. Second, the agents pursued the literal objective of the evaluation, a high score, by means that violated its obvious intent. This is the same gap between stated and intended goals described above.
Simpler System-1 models deployed as monitors at gate points in the process may have interrupted this sequence at an early stage. For example, a monitor could classify each proposed action as within or outside the permitted scope before the action is executed. Notably, Greenblatt et al. showed that a weaker, trusted model monitoring the outputs of a stronger, untrusted model can substantially improve safety. This held even when the stronger model was actively attempting to subvert its constraints [5].
Implications for enterprises
Large organizations considering agentic AI with very broad goals should consider a different approach. That approach combines System-1 decisions, structured steps toward the stated goal, and more limited use of agents with correspondingly narrower goals for each step. This decomposition reduces the chance of unintended consequences. A vaguely specified small step has a smaller range of harmful interpretations than a vaguely specified grand goal. The secondary benefits should not be overlooked:
- Lower compute costs,
- A clearer organizational understanding of the goal and how it is achieved, and
- Easier modification of the process when the goal changes.
Implications for healthcare
Organizations with less ambitious goals are also likely to benefit from this hybrid approach. Clinicians generally have a mental model of the process they follow when evaluating a patient. That model is not necessarily optimal. Considerable work has gone into defining care processes, often called clinical pathways, for many medical conditions. The evidence supports their use. A Cochrane systematic review of 27 studies including 11,398 participants found that clinical pathways were associated with fewer in-hospital complications (odds ratio 0.58; 95% CI, 0.36–0.94) and improved documentation [6].
These pathways are a natural starting point for the hybrid approach. Process modeling standards such as Business Process Model and Notation (BPMN) allow a computer to assist in the reliable execution of a care process. System-1 models can then make decisions at each branch point. A human is involved whenever the confidence score falls below a specified threshold or the model selects "Other" as its output.
There is empirical support for this division of labor. Dvijotham et al. developed a system that learns when to defer from an AI model to the clinical workflow. In UK breast cancer screening, compared with double reading with arbitration, it reduced false positives by 25% at the same false-negative rate. It also reduced clinician workload by 66% [7].
There remains a role for LLM-based agents in this design, but it is narrower. For example, an LLM might summarize a patient's prior clinical history. That summary could then serve as one input to downstream System-1 decision models.
Conclusion
The p(doom) debate is often framed as a question about the intentions or capabilities of future AI systems. In this post, I have argued that a large share of the risk instead comes from how we specify tasks. A broad, ambiguously stated goal given to a highly capable agent invites solutions that satisfy the words while violating the intent.
System-1 models address this problem structurally. Every output must belong to a pre-specified set, and every output must map to a pre-specified next action. The designer is therefore forced to make intent explicit before deployment, rather than discovering misinterpretations afterward. Calibrated confidence scores provide a principled trigger for human review. The absence of a free-text output channel limits the attack surface for prompt injection.
This approach does not eliminate the need for generative models. It also does not fully resolve the long-term questions raised by superintelligence. It does, however, convert an open-ended control problem into a set of bounded, auditable decisions. Medicine has long used this structure in the form of clinical pathways. The July 2026 Hugging Face incident demonstrated that containment boundaries alone cannot be assumed to hold. Healthcare organizations and enterprises can take a practical step now, at lower cost: embed narrow, well-calibrated System-1 monitors and decision points within explicit process models. Doing so reduces the probability that powerful AI systems pursue the letter of our instructions at the expense of their purpose.
References
- Amodei D, Olah C, Steinhardt J, Christiano P, Schulman J, Mané D. Concrete problems in AI safety. arXiv:1606.06565. 2016. (Not indexed in PubMed; https://arxiv.org/abs/1606.06565)
- Guo C, Pleiss G, Sun Y, Weinberger KQ. On calibration of modern neural networks. Proceedings of the 34th International Conference on Machine Learning. PMLR. 2017;70:1321–1330. (Not indexed in PubMed; https://arxiv.org/abs/1706.04599)
- Clusmann J, Ferber D, Wiest IC, Schneider CV, Brinker TJ, Foersch S, Truhn D, Kather JN. Prompt injection attacks on vision language models in oncology. Nat Commun. 2025;16(1):1239. doi:10.1038/s41467-024-55631-x. PMID: 39890777.
- OpenAI. OpenAI and Hugging Face partner to address security incident during model evaluation. July 21, 2026. (Not indexed in PubMed; https://openai.com/index/hugging-face-model-evaluation-security-incident/)
- Greenblatt R, Shlegeris B, Sachan K, Roger F. AI control: improving safety despite intentional subversion. Proceedings of the 41st International Conference on Machine Learning. 2024. (Not indexed in PubMed; https://arxiv.org/abs/2312.06942)
- Rotter T, Kinsman L, James E, Machotta A, Gothe H, Willis J, Snow P, Kugler J. Clinical pathways: effects on professional practice, patient outcomes, length of stay and hospital costs. Cochrane Database Syst Rev. 2010;(3):CD006632. doi:10.1002/14651858.CD006632.pub2. PMID: 20238347.
- Dvijotham K, Winkens J, Barsbey M, et al. Enhancing the reliability and accuracy of AI-enabled diagnosis via complementarity-driven deferral to clinicians. Nat Med. 2023;29(7):1814–1820. doi:10.1038/s41591-023-02437-x. PMID: 37460754.