Can AI hack your software? What the evidence shows about AI cyberattacks
The evidence from 2024 to 2026 says yes. What AI mostly does is make known attacks faster and cheaper, which removes the reason attackers used to skip smaller companies.

Can AI hack software? Yes, and this is documented, not predicted. AI cyberattacks are attacks in which an AI model does part of the attacker’s work: finding vulnerabilities, writing phishing, impersonating people with deepfakes, running reconnaissance or deciding what to steal. AI agents have found real zero-days in SQLite, OpenBSD and FFmpeg, topped a public bug bounty leaderboard, and run most of a live espionage campaign with little human help. The good news: they mostly speed up old attacks, so strong fundamentals still work.
AI cyber threats get plenty of breathless coverage, so every figure here comes from a named 2024 to 2026 source, and forecasts are labeled as forecasts. What engineering leaders need is a clear view of what has been shown to work, how it changes the economics of attacking a mid-sized software company, and what to do next. The short version for smaller companies: size used to predict how much attention attackers paid you. AI is breaking that link.
What AI hacking has already demonstrated
AI systems have already found real vulnerabilities in widely used software: Google’s Big Sleep in SQLite, DARPA’s AIxCC finalists in open source code at about $152 a task, and Anthropic’s Claude Mythos Preview in every major operating system and web browser. XBOW, an autonomous pentester, reached number one on HackerOne’s US leaderboard. The strongest evidence that AI can find exploitable bugs comes from defenders and researchers who published their results. That means the capability is measured and reviewed, not rumored.
Google Big Sleep: an AI agent that found a live SQLite zero-day
Big Sleep is a vulnerability-hunting agent built by Google DeepMind and Google Project Zero. In late 2024 it found a stack buffer underflow in SQLite, one of the most widely deployed pieces of software on the planet, and the bug was fixed before it reached an official release, as The Hacker News reported in November 2024.
Then it got more serious. In mid-2025, Google’s threat intelligence team saw attackers staging a zero-day but could not immediately identify the vulnerability. Big Sleep isolated it: CVE-2025-6965, an SQLite memory corruption bug rated CVSS 7.2, fixed in SQLite 3.50.2 at the end of June and disclosed by Google in July. Google said it believes this was the first time an AI agent had been used to directly foil efforts to exploit a vulnerability in the wild. About three weeks later, Google’s VP of security said Big Sleep had found and reported 20 vulnerabilities, mostly in open source projects such as FFmpeg and ImageMagick.
A defender’s AI won a race against an attacker who already knew about the bug. Both sides of that race can use the same kind of tool.
XBOW: an autonomous pentester at the top of HackerOne
HackerOne runs one of the largest bug bounty platforms, where thousands of human researchers compete to find vulnerabilities in real company systems. By June 2025, XBOW, an autonomous penetration testing system, had reached the number one spot on HackerOne’s US leaderboard. According to XBOW, it had submitted nearly 1,060 reports, all found automatically, with its security team reviewing them before submission to comply with HackerOne’s rules on automated tools.
In the 90 days before its June 2025 write-up, its findings included 54 critical and 242 high-severity vulnerabilities. Across all its submissions to that point, program owners had resolved 130 and triaged 303 more. The point is not that a machine “beat” humans on a leaderboard. It is that an automated system produced a steady stream of valid, serious findings against live production systems owned by real companies, which is exactly the work an attacker wants to automate.
DARPA AIxCC: real bugs found for about $152 a task
DARPA’s AI Cyber Challenge (AIxCC) asked teams to build systems that find and fix vulnerabilities in open source code without human help. At the final competition results announced in August 2025, the systems analyzed more than 54 million lines of code. They found 54 of the 63 synthetic vulnerabilities DARPA had planted (86%) and patched 43 of them. They also found 18 real vulnerabilities that nobody had planted, now being disclosed to the projects’ maintainers, and provided 11 patches for them.
Patches arrived in 45 minutes on average, and the average cost per competition task was about $152. A year earlier, at the semifinal, the systems found 37% of the planted vulnerabilities and patched 25%. For a business leader, $152 is the number to remember: security review used to be priced in weeks of expert time.
Claude Mythos Preview: thousands of zero-days, and a restricted release
In April 2026, Anthropic announced Project Glasswing, which gives a restricted group of partners access to an unreleased model, Claude Mythos Preview, to find and fix vulnerabilities in critical software. The launch partners include AWS, Apple, Cisco, CrowdStrike, Google, Microsoft, NVIDIA, the Linux Foundation and others. Anthropic reported that over the preceding few weeks the model had found thousands of zero-day vulnerabilities, many of them critical, in every major operating system and every major web browser. Examples included a 27-year-old flaw in OpenBSD, a 16-year-old flaw in FFmpeg in code that automated testing tools had run five million times without catching it, and several Linux kernel bugs chained together to go from ordinary user access to full control of the machine.
In its late-May update, Anthropic said it and its roughly 50 partners had found more than 10,000 high- or critical-severity vulnerabilities in about a month, a count that combines its own reports with its partners’. In a separate scan of more than 1,000 open source projects, 90.6% of the 1,752 flagged issues that were checked, mostly by six outside security firms, turned out to be real vulnerabilities. Anthropic also named the new bottleneck: finding flaws had become easier than fixing them, and some open source maintainers had asked it to slow down.
Anthropic has held the model back from public release because nobody yet has safeguards strong enough to stop it being misused. Our own projection (not Anthropic’s claim): capabilities restricted today tend to appear in widely available tools later.
AI-powered cyberattacks already happening in the wild
AI now does most of the work in some real attacks. Anthropic documented a 2025 espionage campaign (GTG-1002) in which AI did 80 to 90% of the work against roughly 30 organizations, and an extortion operation (GTG-2002) that targeted at least 17. Through 2024, the documented criminal use of AI was mostly productivity help: research, troubleshooting code, drafting content. The best public record of what changed comes from AI providers reporting misuse of their own products.
Data extortion run largely by an AI agent
In its August 2025 threat intelligence report, Anthropic described a criminal operation it tracked as GTG-2002. The attacker used Claude Code to automate reconnaissance, harvest credentials and break into networks, then let the AI make tactical and strategic decisions, including which data to steal. The actor targeted at least 17 organizations, across healthcare, emergency services, government and religious institutions. The AI analyzed stolen financial data to decide how much to demand, and ransom demands sometimes exceeded $500,000.
The same report described a criminal with only basic coding skills who used AI to build ransomware and sold packages of it for $400 to $1,200, and North Korean operatives who used AI to get and keep remote jobs at US Fortune 500 technology companies, including passing their technical interviews. Anthropic’s conclusion was that AI has lowered the barriers to sophisticated cybercrime.
The first large-scale attack executed mostly by AI
In November 2025, Anthropic published details of a campaign it attributed with high confidence to a Chinese state-sponsored group, GTG-1002. Detected in mid-September 2025, the operation targeted roughly 30 organizations (tech, finance, chemicals and government) and succeeded in a small number of cases. According to Anthropic’s write-up, the AI performed 80 to 90% of the campaign, with humans stepping in at perhaps four to six critical decision points per hacking campaign. At peak, the AI was making thousands of requests, often several per second.
Anthropic said it believed this was the first documented large-scale cyberattack executed without substantial human intervention. It also noted a limitation: the AI sometimes hallucinated credentials or claimed to have extracted secrets that were actually public. Autonomous attackers make mistakes, just very quickly.
2026: agent swarms and student operators
Anthropic’s September 2026 threat intelligence report goes further. One group, GTG-10007, targeted roughly 50 organizations across education, retail, energy, technology, healthcare, finance and manufacturing, plus several government agencies. Its operators routinely ran “agent swarms,” where a lead AI agent split reconnaissance and post-exploitation work across many sub-agents running in parallel. The group also ran autonomous vulnerability research: its work against a major intrusion-detection product turned up several previously unknown vulnerabilities, and a separate workflow aimed at network appliances produced more than a dozen possible zero-day findings in a single month. Anthropic identified two of the operators as university undergraduates studying computer and communication engineering.
As the report’s first trend puts it: “Sophisticated attacks no longer require sophisticated attackers.”
Malware that asks an AI model for help mid-attack
In November 2025, Google’s threat intelligence group reported the first malware families, such as PROMPTFLUX and PROMPTSTEAL, that query large language models while running and change their behavior mid-execution (PROMPTFLUX was still experimental). Mandiant’s M-Trends 2026 report adds a credential stealer, QUIETVAULT, that checks whether AI command-line tools are installed on a machine and runs its own prompts through them to search for configuration files. Your developers’ AI tools are now part of your attack surface.
The economics flipped: why attacks got cheaper
Three shifts have moved the economics toward attackers: phishing converts better (54% click-through against 12%, per Microsoft), exploitation starts before a patch exists (Mandiant estimates a mean of negative 7 days), and defenders fix more slowly (Verizon puts the median at 43 days for known exploited vulnerabilities). That matters because most businesses are protected less by a wall than by the fact that attacking them costs more than it is worth.
Phishing that gets 4.5x the clicks, and could pay up to 50x more
Microsoft’s Digital Defense Report 2025 reports that AI-automated phishing emails achieved a 54% click-through rate, against 12% for generic phishing emails, a 4.5x increase. A Harvard Kennedy School study cited in the report found the AI-automated emails performed on par with emails written by human experts, and estimated that AI automation could make phishing up to 50 times more profitable against large audiences, by scaling highly targeted attacks to thousands of people at minimal cost. In the incidents Microsoft investigated where it could identify a motive, at least 52% were driven by financial gain and only 4% by espionage alone.
Most attackers want money, and money is available from companies of every size.
IBM’s 2026 Cost of a Data Breach study found that one in four malicious breaches were AI-enabled, a 56% jump in a year, mostly through deepfake impersonation and AI-enabled malware. Those breaches cost about $6 million on average, roughly $1 million more than other malicious breaches.
The exploit window went negative
The gap between a vulnerability becoming known and attackers using it has been shrinking for years. Mandiant’s time-to-exploit analysis put the average at 63 days in 2018 and 2019, 32 days in 2021 and 2022, and just 5 days in 2023. In M-Trends 2026, Mandiant estimates the mean time to exploit at negative 7 days. In other words, exploitation now routinely starts before a patch exists.
In 2025, 28.93% of the known exploited vulnerabilities VulnCheck tracks showed evidence of exploitation on or before the day the CVE was published, according to VulnCheck’s 1H-2026 State of Exploitation report. In the first half of 2026 it was 23.43%: nearly one in four exploited flaws gave defenders no head start.
Mandiant also found that the median time for an initial access partner to hand a compromised network to a second criminal group fell from more than eight hours in 2022 to 22 seconds in 2025.
Defenders are patching slower, not faster
While attackers speed up, many defenders are slowing down. According to Help Net Security’s summary of Verizon’s 2026 Data Breach Investigations Report, exploiting vulnerabilities has become the most common way attackers get in, overtaking stolen credentials for the first time in 19 years. The median time to fully fix vulnerabilities on CISA’s known exploited list rose to 43 days, from 32 the year before. Only 26% of those vulnerabilities were fully fixed by the organizations Verizon studied, down from 38%.
Attackers can exploit a flaw before it is announced. The median organization takes about six weeks to fully fix a known exploited one. That gap is where breaches happen.
Why small companies are targets now
Small companies are targets now because the thing that protected them, the cost of a skilled attacker’s time, is what AI automates. The old reasoning went: “We are a 60-person SaaS company. We are not interesting enough.” That was already weak. AI finishes it off.
Attackers already went where the defenses were thinner. Verizon’s 2025 DBIR found ransomware in 88% of breaches at small and medium-sized businesses, compared with 39% at larger organizations. Third-party involvement doubled to 30% of breaches in the 2025 report and reached 48% in the 2026 one. If you sell to bigger companies, you are their supply chain, and often the easier way in.
AI removes the reason to skip you. Attackers prioritized because skilled human hours were scarce. When an agent swarm handles reconnaissance and initial exploitation, as in the GTG-10007 campaign, the cost of one more target falls toward zero. You no longer need to be worth a skilled human’s week, only the cost of some compute.
The documented targets were not all giants. The AI-run extortion group (GTG-2002) targeted religious institutions and emergency services. The agent-swarm group (GTG-10007) hit education and retail. They were reachable, and that was enough.
Your code is findable. The AI-driven discovery behind Project Glasswing’s 10,000+ serious findings is also shrinking the time between “nobody knows about this bug” and “everyone does.” A vulnerable open source library is a target whether you have 20 employees or 20,000.
The UK’s National Cyber Security Centre put it well in its May 2025 assessment of AI’s impact on cyber threats through 2027. It warns of a “digital divide” between systems keeping pace with AI-enabled threats and a large proportion that are more vulnerable. It expects AI to increase attacks on systems that have not been updated with security fixes, and warns that developers racing to ship AI models and apps may put release speed ahead of security. Plenty of small and mid-sized software companies are in both groups.
Demonstrated vs. projected: a reality check
Coverage of AI cyber threats often mixes what has happened with what might. Here is how the sources above separate.
| Claim | Status | Evidence |
|---|---|---|
| AI can find previously unknown, exploitable bugs in widely used software | Demonstrated | Big Sleep (SQLite, 2024 and 2025), AIxCC (18 real vulnerabilities), Mythos Preview (thousands of zero-days, 90.6% of a checked sample confirmed) |
| AI can find valid, high-severity bugs in live production systems at scale | Demonstrated | XBOW: 54 critical and 242 high findings in 90 days on HackerOne |
| AI can run most of a real intrusion campaign | Demonstrated | GTG-1002: 80 to 90% of the work done by AI across about 30 targets |
| AI lowers the skill needed to run damaging attacks | Demonstrated | GTG-2002 extortion, ransomware sold by a low-skill actor, two GTG-10007 operators identified as undergraduates |
| AI makes phishing more effective | Demonstrated | Microsoft: 54% vs. 12% click-through |
| AI makes phishing up to 50x more profitable | Estimate | A Harvard Kennedy School estimate cited by Microsoft, not a measured figure |
| The time from disclosure to exploitation will shrink further | Projection | NCSC: AI will “almost certainly” reduce it further by 2027 |
| AI-discovered bugs are exploited more often than others | Not supported so far | VulnCheck: only 1.3% of 1,061 AI-attributed vulnerabilities confirmed exploited |
| AI invents entirely new classes of attack | Not seen yet | Google threat intelligence (September 2026): AI layered onto known tradecraft, no fully autonomous attack pipelines seen in the wild |
The demonstrated capabilities are serious. The scariest version of the story, where AI invents unstoppable new attacks, is not what the evidence shows.
The reassuring part: AI amplifies old attacks
Look at what the documented AI attacks actually did: reconnaissance, credential harvesting, exploiting known weakness classes, moving through networks, stealing data, writing phishing emails. None of that is new.
Google’s threat intelligence group said this directly in its January 2025 analysis of how threat actors used Gemini: it did “not see indications of them developing novel capabilities,” and generative AI “allows threat actors to move faster and at higher volume.” That has moved since. Its November 2025 update reported adversaries “deploying novel AI-enabled malware in active operations,” the mid-execution malware described above, and by May 2026 it had identified a zero-day exploit it believes was developed with AI. Its September 2026 update still describes a gradual maturing of tradecraft with AI capabilities layered on, and says it has not yet seen fully autonomous attack pipelines used against targets in the wild. From its frontline incident response work, Mandiant says in M-Trends 2026 that the vast majority of successful intrusions still come from basic human and system failures, and that it does not consider 2025 the year breaches were the direct result of AI.
VulnCheck’s July 2026 report adds a useful counterweight to the doom narrative. Of 1,061 vulnerabilities attributed to AI-assisted discovery, only 14 (1.3%) had been confirmed as exploited in the wild, roughly the rate for all vulnerabilities in the first half of 2026. VulnCheck says its data does not suggest AI-discovered vulnerabilities are inherently more likely to be exploited than ones found by traditional methods. In its words, AI “appears to increase the volume of vulnerabilities that can be discovered, giving defenders an opportunity to identify and remediate them before attackers do.”
That is our view. AI is a speed and volume multiplier on attacks you already know how to stop. Knowing how was never the hard part; doing it consistently and quickly was. Fundamentals you could put off for a quarter now need to happen in days.
Defenders get the multiplier too. IBM’s 2025 Cost of a Data Breach report found that organizations using security AI and automation extensively saved about $1.9 million per breach compared with those using none. In 2026 the saving was almost $2 million, yet among the roughly half of breached organizations running AI agents in security operations, only 18% used them for vulnerability management.
How to defend against AI cyber threats
If AI mostly accelerates known attacks, the defense is to do the known things faster, and to use the same tools attackers use. Here is where we would start with a small or mid-sized software team.
1. Treat patch speed as a core metric
With the mean time to exploit estimated at negative 7 days and a median of 43 days to fully fix a known exploited vulnerability, your patch speed is your exposure window. Track it like uptime. Set targets by severity (days, not weeks, for critical internet-facing issues), automate dependency updates, give the queue an owner, and watch CISA’s known exploited vulnerabilities list.
2. Know exactly what you run
You cannot patch a library you do not know you ship. Keep a software bill of materials, map internet-facing assets, and remove what you no longer use. Attackers’ agents map your perimeter automatically; your map should be better.
3. Close the identity gaps that make automation profitable
AI phishing and deepfakes target people and credentials. Phishing-resistant MFA, least-privilege access and short-lived credentials limit what a stolen password can reach. In Verizon’s 2026 report, only about 23% of third parties fully fixed missing or weak MFA on their cloud accounts, so ask vendors too.
4. Secure your own AI features and developer tools
IBM’s 2026 study found that more than 20% of the breached organizations it studied had a breach involving their AI models or applications, up from 13% a year earlier. In the 2025 study, 97% of organizations with an AI-related breach lacked proper AI access controls (92% in 2026). If you ship AI agents or connect models to internal tools through protocols like MCP, treat prompt injection, data poisoning, model access and tool permissions as real attack surface. Our explainer on agentic AI covers where agents get their permissions and why that matters.
5. Put AI on the defensive side of your pipeline
AI-assisted code review, automated security testing in CI and AI triage of scanner results are available to defenders now. Run them on every pull request, not once a year. At AIxCC prices, continuous testing is no longer only for big companies. Build security checks into your QA and test automation so they run with every release.
6. Plan for the breach you do not prevent
Mandiant calls modern ransomware a choice between paying and rebuilding. Keep offline, tested backups with separate credentials, and rehearse recovery. In Verizon’s 2026 report, 69% of ransomware victims did not pay; the ones who can refuse are the ones who can rebuild.
7. Test like an attacker, and turn findings into fixes
A penetration test only helps if the findings get fixed, the same bottleneck Anthropic flagged in Project Glasswing. Prioritize by real exploitability, fix in code, and re-test.
None of these steps is new. The playbook has not changed; the time you have to run it has.
Where unicrew fits: closing what AI finds
AI has made finding vulnerabilities cheap and fast. Closing them in a live product without breaking it is still engineering work, and it is what unicrew’s cybersecurity consulting is built around: fundamentals, done quickly, and kept up as the product changes.
The usual starting point is a fixed-scope review, with its scope, timeframe and price settled before work starts. We assess the application, its environment and the process that changes both, and return every finding in ranked order, weighing what it would take to exploit against how much it matters to the review or regulation in front of you. You decide who fixes what. A finding is only worth what gets closed, so our engineers can take the fixes too, in the repository and tracker you already use, and re-test what they close. Because a product keeps changing after the assessment, we can also stay attached to delivery, with design review on new features and dependency and pipeline checks inside the build.
The work covers security architecture review, vulnerability management, DevSecOps, penetration testing, AI application security review and incident response readiness, in the stacks we build in (.NET, PHP, Node.js, Angular, React and Vue). It sits inside your product: the application, its environment and the pipeline that ships it. unicrew runs ISO 27001:2022 itself, audited by Quay Audit UK, so the controls we recommend are ones we get examined on. If you sell to bigger companies and their security review is holding up a deal, the Security and Compliance Readiness Sprint is a fixed-scope way to start.
Key takeaways
- AI hacking is demonstrated, not theoretical. Big Sleep, XBOW, DARPA’s AIxCC and Claude Mythos Preview have all found real vulnerabilities in widely used software, and in one case stopped an attack that was already being prepared.
- AI-run attacks are documented. Anthropic has published cases where AI did 80 to 90% of an intrusion campaign, ran data extortion against at least 17 organizations, and coordinated “agent swarms” against roughly 50 targets.
- The economics favor attackers right now. AI-automated phishing got 4.5x the click-through rate of generic phishing, exploitation now often starts before a patch exists (Mandiant: negative 7 days), and the median time to fix a known exploited vulnerability grew to 43 days.
- Being small no longer protects you. Ransomware appeared in 88% of SMB breaches in Verizon’s 2025 data, and AI makes adding one more target almost free.
- AI mostly amplifies old attacks rather than inventing new ones. Patch speed, asset inventory, strong identity controls, secure AI features, AI-assisted testing and tested backups still work. You just need to do them faster.
Frequently asked questions
Yes. Google's Big Sleep found an exploitable SQLite bug in 2024 and in 2025 identified CVE-2025-6965 before attackers could use it. DARPA's AIxCC systems found 18 real vulnerabilities, and Anthropic's Claude Mythos Preview found thousands of zero-days across major operating systems and browsers. Attackers are doing the same: Anthropic reported a group whose AI research produced more than a dozen possible zero-days in one month.
Yes, and in some ways more than before. Verizon's 2025 DBIR found ransomware in 88% of breaches at small and medium-sized businesses, against 39% at large ones. AI automation lowers the cost of each additional target, and documented AI-driven campaigns targeted organizations such as religious institutions, education providers and retailers, not only large enterprises.
Not by itself. A certification like ISO 27001 shows that your security processes passed an audit, but attackers exploit whatever is unpatched today. Verizon's 2026 DBIR found only 26% of known exploited vulnerabilities were fully fixed, while Mandiant now puts the mean time to exploit at 7 days before a patch arrives. Compliance is a floor. Patch speed and continuous testing are what close the gap.
Mostly, but not entirely yet. In the GTG-1002 campaign Anthropic documented in 2025, AI performed 80 to 90% of the work, with humans stepping in at perhaps four to six critical decision points per hacking campaign. The AI also made mistakes, such as hallucinating credentials, so human oversight still mattered to the attackers.
Use AI-assisted code review and automated security testing on every pull request, AI triage of vulnerability scanner results, and automated dependency patching. IBM's 2026 study found that organizations using security AI and automation extensively cut breach costs by almost $2 million on average compared with those using none, yet only 18% of the breached organizations running security AI agents use them for vulnerability management. Pair these tools with human review so findings turn into fixes.


