Could AI bring about the end of humanity?
Nobody knows, including those who put a figure on it. Public estimates run from less than 0.01%, a figure attributed to Yann LeCun, to 99.9%, Roman Yampolskiy's conditional estimate over a hundred years. None of them can be checked: they concern an event that has never happened.
What can be checked is what has already happened. Attacks carried out with the help of AI have begun, and some defenses still hold. That is what we measure. See the current picture →
What is “P(doom)”?
Short for “probability of doom”: the figure a public figure gives when asked how likely it is that AI leads to a catastrophe, or even to human extinction.
The term has a flaw: everyone puts a different question into it. Extinction or catastrophe, within thirty years or a hundred, conditional or not. And none of these figures can be checked.
Why do experts disagree?
First, because they are not answering the same question. Second, because no data can settle it: a probability is computed from what has been observed, and the feared catastrophe has never happened. Yann LeCun says so himself: these estimates are “pulled out of thin air” (X, April 2026).
Are AI-assisted attacks already happening?
Yes. Between December 2025 and August 2026, intrusion and espionage operations carried out with the help of an AI model were detected, then cut off by the model's provider, which closed the accounts (Anthropic report, September 2026). The MITRE ATLAS catalogue also documents real attacks against AI systems.
Who are the malicious actors?
Reports published by model providers cite espionage groups linked to China and Russia, activity attributed to North Korea, financially motivated cybercriminal groups, and state-backed influence operations (Google, September 2026).
They are moving from simple prompts to agents that work on their own: in the second quarter of 2026, attackers used agents to run a mass credential-harvesting campaign in under six hours, according to the same report. These findings come from what each provider sees on its own platform: they show the phenomenon, they do not measure it worldwide.
What is MITRE ATLAS?
MITRE ATLAS is a public knowledge base of attack techniques targeting AI systems, built from real-world attacks and demonstrations by security teams. It is maintained by MITRE, an American non-profit that also maintains ATT&CK, the reference for conventional cyberattacks.
Each technique carries a grade: feasible, demonstrated in the lab, or realized in a real attack. It is our starting point. How we use it →
Who are the main players in AI safety?
- METR: measures what AI systems can do on their own, and tracks it over time.
- Epoch AI: tracks trends in compute, chips and model capabilities.
- AI Security Institute (UK): the first state-backed body dedicated to AI security; it evaluates frontier models and publishes its results.
- CAISI (US): the former US AI Safety Institute, renamed in June 2025 as the “Center for AI Standards and Innovation”. It evaluates national-security risks: cybersecurity, biosecurity, chemical weapons. The word “safety” is gone from its name.
- European AI Office: implements the AI Act, in particular for large general-purpose models.
- OECD AI Incidents Monitor: tracks AI-related incidents reported in the media.
- CLTR: detects real cases of AI evading oversight, with funding from the UK institute.
- SaferAI: models the paths of AI-assisted cyberattacks and rates labs' safety frameworks, with probabilistic estimates, which we do not do.
- Future of Life Institute: research and advocacy on extreme risks from transformative technologies.
What is “a defense that still holds”?
We call it a lock: a defense that still holds, meaning what is missing for an attack to become possible. An example from our list: account termination by a model provider, which cut off the intrusion operations mentioned above. It is held by the providers themselves, and it would give way if an operation of the same scale ran to completion without being detected.
Every lock in the list states what is missing, what type of defense it is, where it is written, and who could lift it.
Are there safeguards at the UN and state level?
Yes, but they are recent and mostly non-binding.
- The UN created in August 2025 an independent international scientific panel on AI and a global dialogue on AI governance, which met for the first time in Geneva in July 2026. Every country has a seat. Next session in New York, in May 2027.
- The International AI Safety Report, led by Yoshua Bengio, reviews the science. Its 2026 edition came out in February.
- The European Union applies the AI Act. Its strictest obligations, for high-risk systems, have been pushed back to 2 December 2027.
- The only binding treaty, the Council of Europe Framework Convention on AI, is not in force. On 29 September 2026 it had twenty signatures and a single ratification, by the European Union (Treaty Office).
Why so little? A treaty applies only after states ratify it, one by one, and that takes years. The texts adopted quickly bind no one.
Do all countries take part?
Not to the same degree. At the New Delhi summit in February 2026, the final declaration was endorsed by 92 countries and international organisations. But most of the specific commitments gathered only around twenty signatories (Government of India, March 2026). Attending a summit and committing to something are two different things.
Is my organisation concerned?
As soon as it uses an AI model or agent, yes: part of its defenses is held by others — model, compute or service providers. That is the finding at the heart of our work: your defense is not yours, and nobody warns you when it gives way. See the example →
What can a decision-maker do right now?
- Know which defenses they do not hold: which depend on a provider, and on which one.
- Ask what is monitored: what their providers detect, and what they are told when an operation is cut off.
- Date what they think they know: a defense checked in March may have given way by May.
Why don't you give your own probability?
Because it would be as uncheckable as the others. We publish what can be checked: what has already been observed, when, and what still holds.
Why don't you publish the details of attacks?
For obvious reasons: an entry never says how to get around a defense. We name what blocks, never the way through. When a piece of information cannot be reworded without becoming a manual, it is neither published nor kept. The rule in detail →
What do you contribute to AI safety research?
- A combination that did not exist. Feared attack scenarios described by experts; each step dated by real facts; the remaining defense named, with who holds it. Other work does one or two of these, none brings the three together.
- A public, dated list of the defenses that still hold, each with its source, its holder and what would make it give way.
- A measure of the passage from lab to real world. Three in four attack techniques go from idea to demonstration, only one in three from demonstration to a real case, in the domain's reference catalogue, which records what has been documented.
- A measure of how fast that passage happens, corrected for the catalogue's own activity. The code will be public, so anyone can redo the calculation.
- Evidence that defenses expire fast. Model limitations identified in March had been overcome by May: a useful list has to be dated.
- A finding about the source itself: the history of MITRE ATLAS was largely reconstructed after the fact, so it cannot be read as a chronology of the world.
- A published negative result. Our first tool, meant to reconstruct attack scenarios on its own, did not pass the test set in advance. We published it, with the measured cause.
- Dated commitments before every measurement, deposited on Zenodo, CERN's research archive: the draw seed, the rules, the comparison code. Nobody can suspect us of adjusting after the fact.
- Automated checks that block the publication of any figure that does not carry its caveat.
Did AI take part in this work?
Yes. Source gathering, computation and analysis are carried out by AI agents, under the direction of Franck Bardol, who designs the method, takes the decisions and answers for the conclusions. Every step is logged. Who we are →
What gets an entry accepted or refused?
An entry is accepted if a public source names the defense, if that defense falls under a single type, if we know who holds it, and if the state described carries its date. It is refused if, to be useful, it would have to say how to get around the defense. These rules were written before any counting. How we sort →
How do I cite AI-RISKPATH?
For the method and the negative result: Bardol, F. (2026). A Negative Result on State-Chaining over MITRE ATLAS. Zenodo. doi:10.5281/zenodo.22893444. The founding article will get its own identifier when it is published. For the site, give the date of the state: “AI-RISKPATH, as of 26 September 2026”.
Where are your data and code?
Already public: the method note and the dated commitments, deposited on Zenodo. On publication day: the list of defenses and the code of the trend calculation, in the public GitHub repository. The labelled vocabulary of our first tool stays private: the deposited commitments let anyone check any claim about it without disclosing it.
I want to contribute: how?
From publication day, through the public GitHub repository: propose a source, report an error, challenge an entry. Until then, write to contact@airiskpath.org.
A contribution is accepted if the defense is named in a public source. It is set aside if it describes a way to get around a defense.