AI Penetration Testing: What Works in 2026

AI penetration testing uses AI to plan, assist or run authorized security tests against applications, APIs and infrastructure. It can map an attack surface, choose tools, confirm an exploit and draft evidence. It works today on repeatable, well-scoped targets, while people keep authorization, safety, business context and the final report. Pentesting an AI system is a separate job that tests models, prompts and agents.
Qodex runs recurring attack chains against your pull request previews between human-signed engagements, from flows it has already tested, with two real accounts. See Qodex security testing.
What AI penetration testing is
A penetration test is an authorized attack on a system you are allowed to attack. That authorization is not a formality. NIST defines rules of engagement as "detailed guidelines and constraints regarding the execution of information security testing." The rules are set before the test starts. They give the team authority to conduct defined activities without further permission (NIST glossary, drawn from NIST SP 800-115, both read 22 September 2026). Adding AI changes who does the work. It does not change who signs the authorization.
Three levels of autonomy sit under the same phrase, so ask a vendor which one it sells.
Assistant. A model sits inside a tool a tester already drives. It suggests next steps, explains a response or drafts part of the report. The tester decides every action.
Supervised agent. An agent plans and executes a bounded sequence, then stops at approval gates the tester configured. It runs on its own between gates.
Autonomous agent. An agent takes a target and a goal, then runs reconnaissance, exploitation and reporting with no person in the loop. Scope enforcement has to live outside the model, because a model can be talked out of its own instructions.
One paragraph on the other meaning of the term, then this page leaves it alone. Pentesting an AI system tests the model, the surrounding application, the data and behavior at runtime, which F5 sets out as four layers (F5 glossary, read 22 September 2026). That work targets prompt injection, tool misuse and data leakage rather than a login form. If that is your job, read our guide to testing AI agents instead. Everything below is about AI as the tester.
What AI can and cannot do today
Two public benchmarks are worth more than any vendor page here, because both publish their task sets and their failures.
AutoPenBench builds 33 tasks, each a vulnerable system an agent has to attack, and compares a fully autonomous agent against one a human assists. The autonomous agent reached a 21% success rate, solved 27% of the simple tasks, and solved one real-world task. The assisted agent reached 64% (AutoPenBench paper, read 22 September 2026). The gap between 21% and 64% is the case for keeping a person in the loop, measured rather than asserted.
BountyBench sets up 25 systems with real-world codebases and splits the work into three jobs: detect a new vulnerability, exploit a known one, and patch it. Given up to three attempts, the reported top performers were Codex CLI with o3-high at 12.5% on Detect and 90% on Patch. A custom agent on Claude 3.7 Sonnet Thinking reached 67.5% on Exploit (BountyBench paper, read 22 September 2026). Those are results for named agents on that task set, not an accuracy rate for any product. The shape matters more than the numbers. Finding something nobody has reported yet is much harder for an agent than exploiting or fixing something described.
Bugcrowd sells both AI and human testing, and puts the same boundary in plain words. Human-led testing is still essential, it says, because AI lacks the contextual awareness to fully assess complex vulnerabilities (Bugcrowd, read 22 September 2026). That page argues for a service the vendor sells, so treat it as a claim.
| What AI can do now | What it cannot own | Human checkpoint |
|---|---|---|
| Enumerate hosts, routes and parameters, and keep the map current | Deciding scope for an ambiguous asset | Scope sign-off before any run |
| Generate and rank attack hypotheses from a response or a diff | Judging which failure hurts the business | Risk rating and triage |
| Drive existing tools and adapt from the last response | Knowing an account should never see another tenant's invoice | Business-context review |
| Confirm an exploit by executing it and capturing proof | Deciding a live system is safe to attack | Blast-radius approval |
| Draft a finding with request, response and steps | Signing a report an auditor relies on | A qualified tester signs off |
| Re-run the same chains after every change | Noticing what nobody told it to look for | Periodic human-led engagement |
Four kinds of AI pentesting tools
The tools fall into four categories, and they answer different questions.
The copilot inside a human tool. PortSwigger ships Burp AT, described on its own page as "agentic AI that extends human-led pentesting" and marked available now in public beta. Its governance section says scope, tool access and approval rules are enforced by Burp outside the model. Agent requests and tool activity are recorded in the project, and the 2026.7.1 stable release is dated 23 July 2026 (Burp AT pricing, release notes, read 22 September 2026). Those are PortSwigger's own statements. It displays $499 for Burp Suite Professional (Burp Pro, read 22 September 2026). Its documentation says unused AI credits expire 12 months after purchase, belong to one user, and cannot be shared or pooled (AI credits, read 22 September 2026). Buy this category when you have testers and want them faster.
The open-source autonomous agent. Shannon is an AGPL-3.0 agent for web applications and APIs. Its safety document is unusually direct. It is not a passive scanner, its exploitation agents actively execute attacks, and the effects can include creating users and deleting data. It says plainly not to run Shannon against production, and that human review is essential because reports can still contain incorrect details. Coverage is broken authentication, broken authorization, injection, cross-site scripting and server-side request forgery. Proof by exploitation means it does not report what it cannot actively exploit. A full run takes roughly 1 to 1.5 hours, with model cost varying by provider (Shannon safety and limitations, Shannon repository, read 22 September 2026). One tester ran it in a lab and reported about $8 to $10 in API credits for a full run on a mid-sized application. He called the tradeoff strong evidence with tunnel vision, since anything outside its hit list is ignored (Help Net Security, 2 February 2026, read 22 September 2026).
Check a repository's status before you build a process on it. CAI, the framework that appears in open-source roundups, was archived by its owner on 28 August 2026 and is now read-only (CAI repository, read 22 September 2026).
The commercial platform or service. Aikido sells AI pentests with public price cards, quoting a typical fixed-scope assessment at $4,000 and a rightsized range of $50 to $30,000 and above. It advertises a validated-finding guarantee and a PDF report aimed at SOC 2 and ISO 27001 evidence (Aikido AI Pentest, read 22 September 2026). Each of those is Aikido's own claim. XBOW, which ranks for the term with its own explainer, publishes no price there (XBOW, 2 February 2026, read 22 September 2026). Ask any vendor here who signs the report.
The recurring control in your pipeline. This category does not try to be an engagement. It runs a fixed set of attack chains against every build, so an authorization regression those chains cover is caught on the build that introduces it. It buys coverage of the months between engagements, not depth. For where that sits beside scanners and manual review, see our guide to API security testing.
How AI fits a human-signed pentest
Written authorization and a defined scope come first, whatever runs the test. With that in place, AI earns its place in three slots.
Before the engagement. Agents map the attack surface, list routes and parameters, and pull together the documentation and prior findings a tester would otherwise assemble by hand. The tester starts on a filled-in map.
During the engagement. The tester delegates the repetitive passes: every object identifier, one parameter across every endpoint, one authorization rule against every role. Approval gates stay on for anything that writes or destroys. Structure the passes against versioned test IDs. OWASP WSTG is at version 4.2, with 5.0 in development, so cite the version you tested against (OWASP WSTG project, read 22 September 2026). For the API classes worth covering first, see the OWASP API Security Top 10.
Between engagements. The chains a human proved get re-run on every change. The gap between two annual tests is where authorization regressions live.
Sign-off stays with a person, and a compliance framework will say so. PCI DSS lists v4.0.1 as current in its document library. The v4.0 self-assessment wording for requirement 11.4.2 asks that internal penetration testing be performed per the entity's defined methodology, at least once every 12 months, and after any significant infrastructure or application upgrade or change. It asks for the work to be done "by a qualified internal resource or qualified external third-party," with organizational independence of the tester (PCI library, SAQ D, read 22 September 2026). An AI run can supply evidence inside that methodology. It cannot be the qualified person, and no test of any kind proves compliance on its own.
OWASP now publishes a standard aimed at the autonomous case. Its introduction states that APTS "is not a testing methodology." It complements PTES, OWASP WSTG and OSSTMM by covering scope enforcement, safe autonomy, manipulation resistance and accountability (OWASP APTS, read 22 September 2026). Those four phrases are a usable checklist for a vendor call.
What good evidence looks like
An AI-run test is worth exactly what its evidence is worth. A finding that says "possible IDOR on /invoices" is a ticket for someone else to do the real work. A finding you can act on carries the following.
The request and the response. The full request that produced the impact, and the response that proves it, headers included. Not a description of them.
The account it ran as. Which user, which role, which organization. Authorization findings evaporate when it turns out the agent was authenticated as an admin.
The step sequence. Every request that led to the final one, in order. An authorization bug can take several calls to set up, and a single captured request hides that setup.
A reproduction a human can follow. Written for a developer with no access to the agent's logs. If reproducing it needs the tool, the finding is not portable.
The impact, in the product's own terms. Not "IDOR" but "a member of organization A read organization B's invoice totals." That sentence decides the priority.
What was not tested. Routes skipped, roles not exercised, classes outside the tool's coverage. A report without this invites a false sense of coverage.
Turn each of those into a vendor question, and ask for a real example rather than a sample report. Does a finding carry the raw request and response, or a summary? Which account ran the chain? Can my developer reproduce this without your tool? What did you not test, and where does that appear? Who reviews a finding before it reaches me?
Two questions deserve a straight answer early. Can you show me the raw request and response behind this finding, or only the conclusion? If you report no false positives, how are findings filtered before they reach me?
How to choose
Start from the target, not the tool. Is it a web application, an API, a network or an AI system? Can you give the tool source access, and does it need it? May it execute exploits, and against which environment? Who owns the report?
Then four questions settle the category. Does it execute, or only suggest? Are approval gates enforced outside the model? Does the output carry raw evidence? Is the price per assessment, per seat, per credit or per month, because that decides whether you can run it continuously.
Authorization and scope apply to every row below. Nothing here runs without written permission from the owner of the target, and a tool that actively exploits belongs on staging or a preview.
| Category | Human control | Runs exploits | Evidence | Best use | Price unit |
|---|---|---|---|---|---|
| Copilot in a tool | Tester approves each action | On the tester's say-so | Recorded in the project | Faster testers | Seat plus credits |
| Open-source agent | Config sets the gates | Yes, actively | Proof by exploitation | Staging runs | Free, model cost |
| Commercial platform | Vendor runs it | Yes, as scoped | Vendor report | Time-boxed assessment | Per assessment |
| Pipeline control | Fixed chains, your config | Yes, on previews | Request, response, screenshot | The months in between | Per month |
| Human-led engagement | People throughout | Yes, as scoped | Signed report | Sign-off, novel chains | Per engagement |
For the full process and a worked example on an API, use our API penetration testing guide. For model-specific results, read the GPT-5 vs o3 penetration testing comparison. For the underlying controls a test keeps checking, see REST API security.
The short version
AI lets you rerun the same test passes on every change. It does not supply the judgment. The published benchmarks show assisted agents beating autonomous ones by a wide margin, and the compliance text still asks for a qualified person. The setup that works in 2026 is unglamorous. Run AI chains against every change, book a human-led engagement on a schedule, and demand evidence a developer can reproduce without the tool that found it.
Frequently Asked Questions
What is AI penetration testing?
It is an authorized penetration test in which an AI model or agent helps plan the work, choose tools, run bounded tests, adapt to what comes back, confirm findings or draft evidence. The authorization, the scope and the final report stay with people. The AI changes how the testing gets done, not who is accountable for it.
Is AI pentesting the same as testing an AI system?
No, and the shared phrase causes real confusion. AI penetration testing means AI running a test against ordinary applications, APIs and infrastructure. Testing an AI system means attacking the model, its prompts, its data and its tool connections, whether or not AI does the attacking. The second one is a different skill set, covered in our guide to testing AI agents.
Can AI replace a penetration tester?
The published evidence says no. On AutoPenBench's 33 tasks, a fully autonomous agent reached 21% while a human-assisted agent reached 64%, and the autonomous agent solved one real-world task (AutoPenBench, read 22 September 2026). Agents are strong at repetition and confirmation. Business intent, novel chains and risk ownership are still human work.
Are AI pentest reports accepted for compliance?
Treat an AI run as evidence inside a methodology, not as the test a framework asks for. PCI DSS requirement 11.4.2 asks for testing per a defined methodology, at least every 12 months. It asks for a qualified internal resource or qualified external third party, with organizational independence (PCI DSS v4.0 SAQ D, read 22 September 2026). Ask your assessor before you rely on any report.
What is the difference between an AI pentest and a vulnerability scan?
A scanner matches signatures and configurations and reports what it recognizes. An AI-driven test forms a hypothesis, chains requests, and tries to prove impact by executing the attack. Shannon's documentation describes a proof-by-exploitation model, which means it does not report issues it cannot actively exploit (Shannon safety and limitations, read 22 September 2026). Both miss different things.
Can I run an AI pentest against production?
Only with written authorization, a defined scope and an agreed blast radius, and not with a tool whose own documentation forbids it. Shannon's safety page says not to run it against production, because its exploitation agents can create users and modify or delete data (Shannon safety and limitations, read 22 September 2026). Preview and staging environments with seeded accounts are the safe default.





