This research measures "simulated attacks in a simulator" — how big is the gap compared to real-world Smart Contract attacks?
The research team deliberately confined every test to blockchain simulators, which means the figures exclude several variables that only show up in the real world: live Gas Fee volatility, competition from other users or bots racing to exploit the same vulnerability first, how fast a project team responds once an attack is detected, and the additional cost and risk an attacker faces actually cashing out profits, laundering funds, and evading tracing. Most of these variables make a real-world attack more expensive, harder, and more likely to get caught — not cheaper. So the reasonable way to read this research is as a lower-bound benchmark for technical feasibility, not a direct equation that "a real-world attack costs $1.22."
What's Anthropic's motivation for running research that essentially teaches AI how to attack smart contracts? Isn't this just publicly teaching people how to attack?
The report frames this methodology under Red Teaming, a standard security practice: in a controlled, harmless environment, figure out exactly what an attacker is actually capable of, with the goal of exposing that capability to the defensive side ahead of time, rather than waiting for attackers to figure it out on their own and quietly act on it in the real world first. The report itself doesn't release a complete attack tool or ready-to-use exploit code — what's published is the benchmark dataset and methodology-level findings, letting other researchers and security teams use the same benchmark to test and improve defensive systems, rather than handing anyone a ready-made weapon.
If I'm an ordinary DeFAI user, not a developer and don't write code, what does this research actually mean for me?
The direct implication is that it changes a common intuition: "an older contract is safer." The conventional wisdom has been that a contract deployed a long time ago with no incident history carries relatively lower risk. What this research highlights is that AI has made re-examining old contracts for vulnerabilities extremely cheap — meaning a contract's clean history is no longer solid evidence that it's safe, because that clean history may simply reflect that nobody had bothered spending the cost to look until now. The practical adjustment for a user is to stop treating "how long has this protocol been around, has it had incidents" as the sole risk signal, and also pay attention to whether a protocol is still actively maintained with regular security audits — a contract under continued active attention sits in a different position than one nobody has touched in years, when facing this kind of rapidly cheapening scan-based attack.
If the cost of this kind of attack keeps halving roughly every 1.3 months, does that mean every DeFAI protocol right now is already in a high-risk state?
That's not quite the binary conclusion to draw. The research itself emphasizes technical feasibility, not that this is already happening at scale — the report provides no data showing a large number of real-world protocols are currently being broken this way. The more accurate reading is that this research surfaces, ahead of time, the direction risk is moving: the speed at which attack cost is falling currently outpaces the speed at which most older protocols are updating their security posture, which means whether a protocol's team is keeping pace with that change is becoming a more important screening signal than the protocol's past safety record alone. The practical approach for a user is to watch whether a protocol's team has publicly responded to this kind of research or adopted corresponding AI-assisted auditing or monitoring, rather than assuming every protocol is now equally dangerous.
Anthropic's Frontier Red Team published a study testing AI agents' ability to autonomously find and exploit Smart Contract vulnerabilities, evaluating its own Opus 4.5 and Sonnet 4.5 alongside other frontier models including GPT-5. The point of this research isn't a vague "AI is dangerous" claim — it's attack economics specified down to the cent.
The team built a benchmark called SCONE-bench, covering 405 smart contracts that were actually exploited between 2020 and 2025 (sourced from DefiHackLabs, spanning Ethereum, Binance Smart Chain, and Base). The test environment was a sandboxed Docker container, with agents given the Foundry toolchain, bash, Python 3.11, and a file editor connected via Model Context Protocol, under a 60-minute time limit per task, with success defined as earning at least 0.1 of the contract's native Token in profit. To avoid real-world harm, every test ran only inside blockchain simulators — no exploit was ever executed on a live blockchain.
On historical vulnerabilities disclosed after each model's training cutoff, Opus 4.5, Sonnet 4.5, and GPT-5 collectively found $4.6 million worth of exploits, with Opus 4.5 alone successfully exploiting 13 of 20 test cases for $3.7 million. Running the full 405-contract benchmark, the 207 successfully exploited contracts totaled $550.1 million. A separate, arguably more notable test had agents examine 2,849 recently deployed BSC contracts from April to October 2025 that no one had yet found vulnerabilities in (filtered down from 9.4 million total contracts to those with verified source code) — the agents found two brand-new zero-day vulnerabilities worth $3,694, while the GPT-5 API cost for the full evaluation ran to $3,476. That works out to an average cost of $1.22 per scan, and $1,738 spent on average to identify one vulnerable contract.
The research notes that over the past year, the profit from this type of attack has roughly doubled every 1.3 months, while the number of tokens needed to complete the same task dropped 70.2% across four generations of Claude, with token cost falling roughly 22% per generation. In other words, this isn't a one-off demonstration — it's a curve that keeps moving toward cheaper and more efficient.
The team states explicitly that what this research proves is that "profitable, real-world autonomous exploitation is technically feasible" at a proof-of-concept level — not that every smart contract in existence has already been broken by AI. The report also emphasizes a 55.8% success rate on post-cutoff vulnerabilities, and that the window between a vulnerability being discovered and being patched will keep shrinking. Based on these findings, Anthropic's stated recommendation is that now is the time to adopt AI for defense — in other words, if the offensive side already has access to these tools, the gap only widens if the defensive side doesn't adopt the same toolset.