Where does our information come from?
From public documents, and only from them.
Our starting point is the MITRE ATLAS catalogue. MITRE is an American non-profit organisation, and ATLAS is the domain's reference: it records attack techniques against AI systems and the cases in which they have been observed.
To find out what still holds, we read what is published by those who test these systems: public evaluation institutes such as the UK AI Security Institute, the labs that evaluate their own models, incident reports, and the providers who describe their controls.
How do we sort it?
We look for locks. A lock is a defense that still holds: what is missing for an attack to become possible.
The essential rule: the lock must be written in black and white in a published document. We never guess it. When nobody has checked whether a defense still holds, we say exactly that: not tested. That is information in its own right, and nobody publishes it.
Going further: the eight types of defense
Capability (the model cannot do it yet), resource (money, compute), identity (payment, account, ID), access (privileged rights), duration (lasting without being detected), tacit knowledge (know-how absent from written sources), installed control (someone deliberately maintains it: rate limits, filters, sandboxes) and not tested.
Each lock gets a single type. If two types fit, it is marked not tested rather than settled by guesswork.
What do we refuse to publish?
For obvious reasons, an entry never says how to get around the defense.
Often, rewording is enough. “Identity checks by compute providers block an autonomous agent from opening an account” is a lock. We do not name the provider that skips them: that would be an attack manual.
When rewording is impossible, the entry is neither published nor kept. This happened twice among the thirteen candidates examined for the list of 26 September, from otherwise excellent sources.
What do we not know?
Three things, which we write down rather than keep quiet.
- The catalogue records what has been documented. We cannot yet tell a field that is accelerating from a catalogue that is better kept.
- Public tests are run without defenders. Evaluators' test ranges have no monitoring team and no detection. So nobody knows what the same attack would achieve against a well-defended system.
- What held yesterday may give way tomorrow. Four model limitations published in March had been overcome by May. These are four observations, not a general law, but they are enough: our list gives the state at a date, and each entry carries its own.
Why the list is dated
We first wanted a list that would stay valid over time. That is impossible, and we measured it.
Under a rule written before any counting, only one scenario in ten met the conditions for a durable list. Under a rule also written before any counting, but for the state at a given date, eight candidates out of thirteen meet them. We publish both results side by side, with both rules.
How can you check us?
Everything we claim can be redone by someone else.
- Every lock cites its source.
- Every figure can be recomputed with a single command, from public sources whose version is pinned.
- The article is archived with a permanent identifier (DOI), which dates what we published and prevents it from being rewritten afterwards.
We also publish our failures. Before this list, we tried to build a tool that would automatically reconstruct attack scenarios. The test was set in advance, and the tool did not pass it. We published that negative result, with the means to check it.