On September 9, 2026, Jacob Coxon, a pretraining researcher who had worked at both OpenAI and Anthropic, resigned and posted a thread on X stating that the two labs are "racing straight to self-improving superintelligence and gambling with our lives." The post accumulated more than 170 million views within days. Evan Hubinger, an alignment science lead at Anthropic, responded publicly that Coxon's characterization was correct and that he personally placed the probability of AI-driven human extinction within the next decade at greater than 10% (1). The exchange is controversial and I risk people on both sides immediately discounting this if it doesn’t agree with their view. I ask that you read through because this is an issue that we need to address and get consensus on as a society. This post doesn’t attempt to settle anything, but it does put two forecasts side by side that are routinely invoked in the same conversation about the same technology despite resting on entirely different kinds of evidence: the probability of AI-caused catastrophe, commonly abbreviated p(doom), and the probability of AI-driven benefit, which has no equally established shorthand but which I will refer to as p(boom). This piece examines what each forecast is actually built on, since the two are not evidentiarily comparable in the way casual discussion tends to imply.

P(doom): a construct built on elicitation, not observation

P(doom) estimates are, almost without exception, subjective probability judgments elicited from individuals rather than measurements derived from data. This is not necessarily a flaw — an extinction event by definition can have no track record to calibrate against, so elicitation is close to the only available method — but it does mean the resulting numbers should be read as structured opinion, not as output of a predictive model in the way an actuarial estimate would be.

The dispersion of individual estimates is large. In the most comprehensive published survey, 2,778 AI researchers were asked to estimate the probability that AI would cause outcomes "extinction-level" in severity; the proportion assigning at least a 10% probability ranged from 37.8% to 51.4% depending on question framing, and the shift in these numbers alone illustrates how sensitive elicited probabilities are to how the question is asked (2). Named individuals span nearly the full range: Yann LeCun and Marc Andreessen have placed the figure near zero, while Eliezer Yudkowsky and Roman Yampolskiy have placed it above 95%, generally on the argument that the technical problem of aligning a system's goals with human intent remains unsolved (3). Aggregated forecasting platforms, which pool many individual forecasters and are sometimes treated as a more disciplined synthesis, place the risk of human extinction from any cause by 2100 at approximately 5%, of which roughly 3 percentage points are attributed specifically to advanced AI (4). Hubinger's figure, cited above, sits within this same wide distribution rather than outside it.

One threat vector in this space is measurable rather than elicited. In 2022, researchers at Collaborations Pharmaceuticals inverted a molecule-generation model originally built to penalize toxicity, instead rewarding it, and the model produced 40,000 candidate toxic molecules in under six hours, including structural analogs of the nerve agent VX and compounds predicted to be more toxic still (5). This result is frequently cited in the p(doom) discourse because it is one of the few data points in that discourse that is not an opinion: it is a documented computational output. It is worth noting, as the original authors do, that generating a molecular structure is a substantial distance from synthesizing, weaponizing, and deploying it, so the finding constrains the plausibility of one specific mechanism without constraining the probability of the broader catastrophic scenarios that dominate the survey literature above.

P(boom): a construct built on measurement, with its own caveats

P(boom) claims, by contrast, are typically anchored to outputs that have already been produced and can, in principle, be independently checked, which is a different evidentiary posture even before any single claim is evaluated for how far it generalizes.

In protein design, a diffusion-based generative model fine-tuned from a structure-prediction network was shown to design binding proteins that were subsequently expressed, purified, and validated experimentally, including a binder whose cryo-electron microscopy structure was nearly identical to the computational design (6). Follow-on work extended the same approach to flexible peptide targets and produced binders with picomolar-range affinity directly from computation, without subsequent experimental optimization (7). These are measured outcomes — binding affinities determined by biophysical assay, structures determined by electron microscopy — rather than projections.

Scale claims are more heterogeneous in evidentiary quality. Caris Life Sciences has reported, in regulatory filings associated with its 2025 initial public offering, that its AI-driven molecular profiling platform has processed more than 849,000 cancer cases and generated measurements on more than 38 billion molecular markers (8). These are large, specific, company-reported figures; they have not been subjected to independent peer review in the way the protein-design results above have, and the distinction matters for how much weight the figures should carry. Similarly, AstraZeneca's 2025 collaboration with CSPC Pharmaceuticals for AI-assisted discovery of oral therapeutics is structured as $110 million upfront with up to $1.62 billion in development milestones and up to $3.6 billion in sales milestones, for a total potential value near $5.3 billion (9) — a real capital commitment, but one whose payout is explicitly contingent on results that have not yet occurred. A 2025 survey of 127 pharmaceutical and life-sciences executives found that 93% anticipated increasing their data, digital, and AI investment in the following year (10), which documents industry sentiment and capital allocation intent rather than realized clinical benefit.

The asymmetry, and its limits

The comparison is not, therefore, that p(doom) is baseless and p(boom) is proven. It is that the two forecasts fail differently. P(doom) estimates cannot be empirically calibrated in principle, since the event in question has no precedent and cannot be observed short of its occurrence; the best available evidence is expert judgment, which is legitimate but should be reported as judgment, with its documented sensitivity to framing, rather than as a measured quantity. P(boom) claims can be more closely checked against reality, and several of the ones considered here — the RFdiffusion binder structures and affinities in particular — already have been, but that checkability does not extend evenly across the category. A validated binding affinity is a different kind of evidence than a self-reported case count or a stated investment intention, even though both circulate under the same "AI in biotech" heading. Treating an executive-sentiment survey as though it carried the same evidentiary weight as an electron-microscopy-confirmed structure would be the same scope error, applied to optimistic claims, that afflicts uncritical treatment of p(doom) figures on the pessimistic side.

What follows for the reasonable observer

None of the estimates reviewed above are precise enough to suggest a specific action threshold, so the practical response for a clinician, executive, or researcher is not to adopt a single p(doom) or p(boom) figure and act as though it were settled. The more defensible move is behavioral rather than numerical: under genuine uncertainty with high stakes on both tails, standard practice favors actions that perform acceptably across a wide range of plausible true values over actions that are justified only if one specific estimate happens to be correct. This is a familiar posture in clinical decision-making under diagnostic uncertainty, where a clinician facing a contested probability of a serious but unconfirmed diagnosis generally prefers a low-cost, reversible, information-generating step — an additional test, a period of monitoring — over either committing to an aggressive irreversible intervention or dismissing the possibility outright.

Applied here, this favors investment in verification and evaluation infrastructure largely independent of where one's own p(doom) estimate falls. Red-teaming, external audits, incident reporting, and capability-specific benchmarking are worth their cost if the probability of catastrophic outcomes is high, and remain worth their considerably lower cost even if that probability is closer to the low single digits many capabilities researchers report (2). A position staked entirely on either the accelerationist or the decelerationist end of the distribution, by contrast, is only correct if that specific estimate turns out to be right — a considerably weaker basis for action given how much the surveyed estimates move with question framing alone.

For the narrower and more immediate question of whether to adopt or restrict a specific AI capability, the same scoping discipline argued for in the construct above applies with equal force to risk and benefit claims. A clinician deciding whether to deploy an ambient documentation tool is not meaningfully informed by an aggregate estimate of civilizational risk, any more than a laboratory deciding whether to synthesize a computationally designed protein binder is informed by industry-wide investment sentiment. Each of those decisions is answerable, if at all, by evidence at its own scale — the verified error profile of the specific tool, the specific binder's measured affinity and specificity — and treating either the pessimistic or the optimistic aggregate figure as decision-relevant at that scale is the same category error in both directions.

Conclusion to Part 1

Coxon's resignation and Hubinger's response occurred in the same week that AI-assisted protein design continued to produce independently verifiable structures and pharmaceutical companies continued to commit real capital to AI-driven discovery programs. Both facts are true, and neither cancels the other, because they are not answers to the same question. Before treating any p(doom) or p(boom) figure as evidence for a broader claim about the technology, it is worth asking what kind of claim it actually is: an elicited judgment, a validated measurement, or a stated intention. The three are not interchangeable, and the discourse around advanced AI would be better served by keeping them separate than by aggregating them into a single number that neither side actually possesses.

I am calling this part 1, because this story is going to develop and as new facts are discvovered, we must all update our assessment of the relative risks. More to come…


References

  1. Hubinger E [@evanhub]. X post responding to Jacob Coxon's resignation thread. September 9, 2026. (Social media post; not a journal article; not indexed in PubMed; https://x.com/hilbertspaess/status/2097476196791709843)
  2. Grace K, Stein-Perlman Z, Weinstein-Raun B, et al. Thousands of AI authors on the future of AI. arXiv. 2024. (Not indexed in PubMed; https://arxiv.org/abs/2401.02843)
  3. AI Impacts. Surveys of expert opinion on AI existential risk. (Wiki/compiled source, not a journal article; not indexed in PubMed; https://wiki.aiimpacts.org/uncategorized/ai_risk_surveys)
  4. Metaculus. Aggregate community forecast, probability of human extinction by 2100, as cited in: The economics of p(doom): scenarios of existential risk and economic growth in the age of transformative AI. arXiv. 2025. (Not indexed in PubMed; https://arxiv.org/abs/2503.07341)
  5. Urbina F, Lentzos F, Invernizzi C, Ekins S. Dual use of artificial-intelligence-powered drug discovery. Nat Mach Intell. 2022;4(3):189-191.
  6. Watson JL, Juergens D, Bennett NR, et al. De novo design of protein structure and function with RFdiffusion. Nature. 2023;620(7976):1089-1100.
  7. Vázquez Torres S, Leung PJY, Venkatesh P, et al. De novo design of high-affinity binders of bioactive helical peptides. Nature. 2024;626(7998):435-442.
  8. Caris Life Sciences, Inc. Registration statement (Form S-1) and related IPO disclosures. 2025. (Company regulatory filing; not a journal article; not indexed in PubMed; https://medcitynews.com/2025/06/caris-ipo-precision-medicine-cancer-detection-companion-diagnostic-techbio-ai-cai/)
  9. AstraZeneca. AstraZeneca enters strategic collaboration with CSPC Pharmaceuticals focused on AI-enabled research. Press release. June 13, 2025. (Not a journal article; not indexed in PubMed; https://www.pharmtech.com/view/astrazeneca-partners-on-ai-enabled-research-with-cspc-pharmaceuticals)
  10. ZS Associates. 2025 AI trends: life sciences leaders on data, digital and AI. 2025. (Industry survey; not a journal article; not indexed in PubMed; https://www.zs.com/insights/2025-survey-data-digital-ai)