Four AI models scored Attack the Glass against fifteen of the most serious security threats on the record, working separately and without seeing each other’s answers. Each model answered the same ten questions about every threat on a nought to ten scale, and calibrated on Log4Shell, Stuxnet and phishing before scoring anything else. ATG came first of sixteen on systemic severity, the composite that measures how hard a threat is to see, trace, outlast, fix and catalogue. Three of the four runs put it first outright, and the fourth put it second. Their four scores span 1.20 points, closer agreement than the models reached on eight of the fifteen benchmarks. ATG is also the only subject on the board with no question scoring below 6.
7.94out of 10
Systemic severity
1st of sixteen, ahead of Spectre and Meltdown at 7.45
010
Four runs: 7.28, 8.00, 8.00, 8.48
The exercise exists because a severity claim made by the people who found a vulnerability is worth very little on its own. Four independent readers applying the same fixed questions to the same sixteen subjects is worth more, and it is checkable.
The ten questions
Each model received the same ten questions about every threat, scored zero to ten, with written anchors at each end so the scale meant the same thing to all four. Higher always means worse for the defender. Two of the questions are inverted: a threat that is cheap to reach and hard to fix scores high, not low.
How much of the world’s computing does it reach?
How cheap is it to pull off?
How precisely can it be aimed?
How hard is it to spot while it is running?
How hard is it to trace afterward?
How long can someone sit inside it before anyone notices?
How bad is the worst case?
Can it actually be fixed?
Do our standards and catalogues have anywhere to put it?
How much confirmed real-world use is on the public record?
Every model calibrated on Log4Shell, Stuxnet and phishing before scoring anything else, in that order, because those three span the range. ATG was scored last, after the scale had settled.
Two composites come out of the ten. Systemic severity weights reach, detection difficulty, remediability, dwell, framework blind spot and attribution resistance: how hard a threat is to live with. Acute severity weights impact ceiling, demonstrated use, access cost and targeting precision: how hard it hits. They share no question between them.
The result
ATG averages 7.94 on systemic severity, first of sixteen. Spectre and Meltdown follow at 7.45, then prompt injection at 6.93. On acute severity ATG averages 7.64, eighth of sixteen, behind phishing, Log4Shell, SQL injection, memory-safety bugs, cross-site scripting, WannaCry and prompt injection.
Where ATG lands against the other fifteen
Every subject sits in one of four quadrants, split at the median of the sixteen rather than at a fixed threshold. ATG shares the severe-on-both-counts quadrant with three others and sits a full point clear of the nearest of them on systemic severity.
The full ranking, with the spread across the four runs
The spread across the four runs matters as much as the average. ATG’s four scores span 1.20 points, tighter than eight of the fifteen benchmarks. XZ Utils spans 2.49, Spectre and Meltdown 2.35, BGP hijacking 2.14. The models agree about ATG more than they agree about several threats the field has known for a decade.
What each of the ten questions measuresEvery threat scored on every question
Reading a threat against the field on a single question is what the matrix supports. ATG is the only subject with no question scoring below 6. The next tightest floor is prompt injection at 5.3, and most benchmarks bottom out between 0.5 and 2.5. Every other threat here has at least one question on which it is simply not a problem.
How each benchmark threat differs from ATG
Measuring each threat against ATG rather than against the scale shows where the differences actually sit. The same two bars run long in almost every panel: remediability and framework blind spot. Log4Shell scores 7.0 and 7.5 below ATG on those two, WannaCry 7.2 and 7.5, SQL injection 6.0 and 7.2. Demonstrated use runs the other way in eleven of the fifteen.
Threat by threat
The fifteen in order of how closely each one’s ten-question shape matches ATG’s. Profile distance is the root-mean-square gap across the ten, so a smaller number means a more similar shape rather than a similar severity.
The rule under each name runs from nought to ten. The filled part is that threat’s own systemic severity and the green line marked ATG stands at 7.94, so a rule that fills past the line belongs to a threat the models rated above ATG. None of them do.
Prompt injection
ATG
Systemic 6.93Distance 1.58
A vulnerability class rather than a single event, named in 2022 once large language models began acting on text handed to them by other people. OWASP ranks it first in its Top 10 for large language model applications, and no general fix has been published.
The closest match to ATG's shape of anything on the board, at a profile distance of 1.58. Both are classes rather than incidents, both sit outside what the catalogues describe, and both are cheap to attempt: prompt injection scores 9.8 on access cost against ATG's 7.2, the largest single gap in its favour. Where it falls away is dwell and framework coverage, 2.8 and 2.0 below ATG. An adversary holding a prompt-injection position does not hold it quietly for years in the way a supply-chain position can be held.
Spectre and Meltdown
ATG
Systemic 7.45Distance 2.40
Disclosed on 3 January 2018. Spectre reaches Intel, AMD, ARM and IBM POWER; Meltdown covers Intel processors going back to 1995 along with IBM POWER and some ARM cores, and the mitigations cost measurable performance on every machine that took them.
Second of sixteen on systemic severity, and the only threat that outranks ATG in any single run. It reaches further, 9.0 against 7.5, because it sits in silicon rather than in a supply chain. It is also the one benchmark that shares ATG's shape of being structurally hard to fix while barely exploited, at 2.5 on confirmed use against ATG's 6.0. The frameworks handle it better, at 3.8 on blind spot against 8.0, because a processor flaw has a product, a vendor and a CVE number.
BGP hijacking
ATG
Systemic 5.45Distance 2.70
Routes on the public internet are accepted largely on the say-so of whoever announces them, and have been since the 1990s. Pakistan Telecom took YouTube off the air for much of the world on 24 February 2008, and an April 2018 hijack of Amazon Route 53 redirected MyEtherWallet users and took about 160,000 dollars in Ether.
Another trust-by-address problem: a route is believed because of where it claims to come from, which is the same mistake ATG exploits one layer up. It is better demonstrated than ATG, 8.2 against 6.0, with a long public record. It is far weaker on dwell and detection, 4.2 and 3.2 below ATG, because a hijacked route is visible to anyone watching the global routing table and is usually withdrawn within hours. Nothing equivalent watches what a browser assembled.
SolarWinds and SUNBURST
ATG
Systemic 6.03Distance 2.71
A trojanised update to SolarWinds Orion shipped to around 18,000 customer organisations from March 2020 and was found in December that year. Fewer than 100 were exploited further, most of them United States federal agencies and technology firms, and Washington named Russia's SVR as responsible in April 2021.
The supply-chain compromise everyone reaches for first, and the models score it level with ATG on detection difficulty at 8.2 and on dwell at 7.8 against ATG's 8.0. It cost far more to build and left far more behind. SolarWinds scores 2.0 on access cost against ATG's 7.2, because it took a state programme to build. It is also 4.0 easier to remediate and 4.8 better covered by frameworks, since a signed build from a named vendor is something the existing controls know how to reason about.
Memory-safety and overflow bugs
ATG
Systemic 5.60Distance 3.09
The oldest class of software flaw still exploited daily, reaching back to the Morris worm in 1988. Microsoft has reported that around 70 per cent of the vulnerabilities it assigns a CVE each year are memory-safety issues, and Chromium reports the same share of its high severity bugs.
Maximally demonstrated at a flat 10 on confirmed use, and level with ATG on damage at 9.0 impact ceiling. Forty years of exploitation produced an entire industry of mitigations, which is why it scores 1.2 on framework blind spot against ATG's 8.0: every catalogue, standard and toolchain has somewhere to put a buffer overflow. It is also 3.0 easier to spot while running. The comparison is instructive rather than flattering: the field solved the visibility problem for memory safety and has not started on this one.
The XZ Utils backdoor
ATG
Systemic 5.39Distance 3.16
A contributor working as Jia Tan spent two years earning maintainer trust on xz, a compression library almost every Linux distribution carries, then shipped a backdoor in versions 5.6.0 and 5.6.1 in February 2024. Andres Freund found it on 29 March while looking into slow SSH logins, by which point it had reached rolling distributions such as Arch and Tumbleweed but not the enterprise releases it was aimed at.
The closest thing to ATG in tradecraft, a long-game maintainer infiltration aimed at a component everything downstream would inherit. It scores marginally above ATG on both detection difficulty and attribution resistance, at 8.5 and 8.0. Two things pull it down. It was caught before it shipped, so confirmed use is 3.0 and reach is 4.0 below ATG, because it never reached the machines it was aimed at. Once found it was fixable, at 2.2 on remediability against ATG's 8.2, because there was a specific commit to revert.
Stuxnet
ATG
Systemic 5.93Distance 3.17
Discovered in June 2010 and built to reach the uranium enrichment plant at Natanz, where analysts estimate it destroyed about a thousand centrifuges. It used four zero-day vulnerabilities and spread to well over a hundred thousand machines worldwide, although the payload fired only on one Siemens controller configuration.
The precision extreme of the whole set, at 9.8 on targeting and a perfect 10 on impact ceiling, both above ATG. It is also the most expensive thing here at 1.0 on access cost against ATG's 7.2, and the narrowest in reach at 1.8 against 7.5. That trade is what the comparison exposes. Stuxnet showed what a manipulated operator screen does to physical plant, and it took a state programme and years to reach a few thousand centrifuges. ATG scores as reaching the same class of effect through an ordinary commercial supply chain.
Phishing and social engineering
ATG
Systemic 6.18Distance 3.20
The oldest technique on this list still running at scale every day, and the one route in that no patch closes. Verizon's annual breach reports put a human element in 60 to 74 per cent of breaches across recent editions, a category wider than phishing that also counts error and stolen credentials.
The highest acute score on the board at 9.23, and the threat ATG most resembles in economics: universal reach at 10.0, free to attempt at 9.8, and no patch that closes it at 8.8 on remediability. Both attack the person rather than the system. They separate on visibility. Phishing scores 5.0 on detection difficulty and 1.5 on framework blind spot against ATG's 8.2 and 8.0, because a phishing email is a message that exists on a mail server and can be reported, blocked and trained against. ATG leaves nothing to report.
Cross-site scripting
ATG
Systemic 4.15Distance 3.97
Named by Microsoft engineers in January 2000 and on the OWASP Top 10 ever since, folded into the injection category in the 2021 edition. Twenty-five years on it is still among the most reported web vulnerability classes.
Named by the models as one of ATG's closest analogues when they were asked in prose, and the resemblance is real: both end with unintended code running in a page at that page's own privilege. The two differ in how the code gets there. Cross-site scripting arrives through unsanitised input and is fixed by sanitising it, which is why it scores 2.2 on remediability against ATG's 8.2 and 0.8 on blind spot against 8.0. It is also exhaustively demonstrated at 10.0. ATG arrives through an authorised dependency the application asked for, so there is no input to sanitise.
SQL injection
ATG
Systemic 3.83Distance 4.06
The earliest widely cited public description ran in Phrack in December 1998, written by Jeff Forristal under the name rain.forest.puppy. Parameterised queries have answered it in principle since the early 2000s, and scanners still find it in production applications every year.
The most catalogued class in the set at 0.8 on framework blind spot, and among the most thoroughly exploited at 10.0. It is cheap and heavily automated, at 9.2 on access cost. Everything that makes it tractable is what ATG lacks: parameterised queries fix it, scanners find it, and every standard names it. It sits 4.5 below ATG on detection difficulty and 5.8 below on remediability.
Heartbleed
ATG
Systemic 4.19Distance 4.07
Disclosed on 7 April 2014, a bug in OpenSSL 1.0.1 through 1.0.1f that let anyone read 64 kilobytes of a server's memory at a time, private keys and session data included. Netcraft estimated that around 17 per cent of the internet's trusted HTTPS servers, roughly half a million of them, were exposed on the day it was announced.
A single defect in a single library that happened to be almost everywhere, at 7.0 reach. It scores respectably on detection difficulty for its time, 6.2, because reading memory left nothing in the logs. Everything after that diverges. There was a patch, so remediability is 1.5 against ATG's 8.2, and the frameworks had a CVE and a version range to work with, so blind spot is 0.5 against 8.0. It also could not be aimed, at 4.0 on targeting against ATG's 8.8.
MOVEit and Cl0p
ATG
Systemic 3.46Distance 4.07
The Cl0p group began exploiting a flaw in Progress MOVEit Transfer on 27 May 2023 and worked through its customers over the following weeks. Emsisoft closed its tally in June 2024 at 2,773 organisations and about 95.8 million individuals, a floor drawn from disclosed breaches rather than a full count.
A mass-exploitation campaign against one managed file transfer product, well demonstrated at 8.5. Its profile inverts ATG's on everything structural: reach 5.0, dwell 3.8, attribution 3.5, remediability 1.5 and blind spot 1.0, all far below. It was a product vulnerability with a vendor, a patch and a victim list, which is precisely the kind of event the response apparatus was built for.
Log4Shell
ATG
Systemic 3.58Distance 4.57
Disclosed on 9 December 2021 in Apache Log4j 2, a logging component bundled into a very large share of Java software, where a crafted string written to a log ran code on the server. CISA put the exposed population at hundreds of millions of devices, and patching ran through the end of that year.
One of the three threats every model scored first to calibrate the scale, and the most instructive contrast in the set. It beats ATG on reach, access cost and confirmed use, at 8.5, 9.5 and 10.0. It loses on everything that determines how long you live with a thing: 3.2 on detection difficulty, 1.2 on remediability and 0.5 on blind spot, against ATG's 8.2, 8.2 and 8.0. Log4Shell was an emergency substantially over within months because there was a version to upgrade to. ATG has no version to upgrade to.
NotPetya
ATG
Systemic 3.10Distance 4.60
Launched on 27 June 2017 through a compromised update to M.E.Doc, Ukrainian tax software, then spread worldwide within hours and wrecked the machines it reached. The White House later estimated total damage at about 10 billion dollars, with Merck at roughly 870 million, FedEx's TNT unit near 400 million and Maersk around 300 million.
The most destructive event on the board by outcome, at 9.5 impact ceiling, above ATG's 9.0. It is also the loudest thing here. Detection difficulty 2.0 and dwell 2.0 against ATG's 8.2 and 8.0, because it destroyed machines immediately and visibly. It scores 3.2 on targeting against ATG's 8.8: it was aimed at one country and hit the world. Impact and stealth are separate axes, and NotPetya is the clearest case of maximum impact with no stealth at all.
WannaCry
ATG
Systemic 2.53Distance 5.37
Began on 12 May 2017 using EternalBlue, an SMB exploit taken from a leaked NSA toolkit, and reached an estimated 200,000 to 300,000 machines across about 150 countries within days. England's NHS was among the worst hit, with up to 70,000 devices affected at a cost later put near 92 million pounds.
The furthest from ATG in shape of anything scored, at a profile distance of 5.37, and last of sixteen on systemic severity. Indiscriminate at 1.0 on targeting, immediately visible at 1.5 on detection, over in days at 1.2 on dwell, and patchable at 1.0 on remediability. It was severe because it spread fast and encrypted things, not because it was hard to see or hard to fix. ATG is the opposite case on all four.
The four runs
Model
Tier
ATG systemic
ATG rank
Claude
Frontier
8.48
1st
Gemini 2.5 Pro
Frontier
8.00
1st
GPT-5.6
Frontier
8.00
1st
Gemini 3.8 Flash
Flash
7.28
2nd
What these numbers do not show
Four limits are worth stating plainly.
ATG was scored with the technical whitepaper supplied. The other fifteen were scored from what the models already knew. That favours ATG on the questions where detail helps, and it is the first thing a careful reader should weigh.
One of the four models is a Flash-tier model, smaller and faster than the other three. It is marked separately on every chart. It produced the most sceptical read of ATG and is the run that ranked it second.
Detection difficulty and confirmed use are not independent. A threat engineered to defeat instrumentation shows little confirmed exploitation partly because of that, so ATG reads low on one question for the same reason it reads high on another. Confirmed use carries no weight in systemic severity, so the systemic ranking does not turn on it.
A fifth run was set aside. It scored ATG at 9.30 systemic with five perfect tens, well outside the pattern of the other four, and it informs none of the numbers here.
What the models read
ATG was scored from the technical whitepaper. A PDF of it sits in thelibrary, along with the technical primer, the playbook and the FAQ. Readers who want plain English can start with theintroduction instead.