Always-on AI pentesting for your web apps and APIsAlways-on AI pentestingStart an AI pentest →
Learn
What Is Agentic Pentesting?
Agentic pentesting is penetration testing done by AI agents that decide their own next step. A coordinator splits the target into testing areas, specialist agents test them in parallel, read each response, change approach when something is blocked and chain requests into real attacks. People set the scope and review the results.
How does an agentic pentest run?
It starts with discovery. The system maps the app, signs in with the accounts you've given it and works out what kind of target it is: a rendered web app, a single-page app or an API.
Then a coordinator agent plans the work. It splits the target into testing areas such as authentication, access control, injection and business logic, and hands each one to a specialist agent. The specialists work in parallel waves. After each wave, the coordinator looks at what came back and reshapes the queue: it follows up on promising leads, re-queues areas that were cut short and drops work that turned out to be pointless.
Inside an area, the agent behaves more like a tester than a script. It sends a request, reads the response and decides what to try next. When a direct request is refused, it tries a different route to the same data. When one weakness gives it a foothold, it chains that into the next request, for example using a leaked ID from one endpoint to ask another endpoint for a different user's record.
None of that is worth much without the last step. Agents make mistakes, so what they report has to be checked against the live app before anyone sees it. We cover why in proof-backed pentesting.
Agentic, AI-assisted or an LLM on a scanner: what's the difference?
All three get sold as "AI pentesting". The question to ask is who decides the next request.
LLM on a scanner
AI-assisted testing
Agentic pentesting
Who picks the next request
The scanner's fixed rules
A human tester
The agent, inside the approved scope
What the AI does
Summarises and ranks output
Suggests payloads and next steps
Plans, attacks, adapts and chains
Multi-step issues
Rarely found
Found if the tester has time
Found when the agent follows the chain
Can repeat on a schedule
Yes
No, needs a person each time
Yes
Proves findings
Usually not
The tester does
Only if the platform adds a proof step
That last row matters more than it looks. Agentic doesn't mean proven. An agent can be confidently wrong, so ask any vendor what happens to a finding between the agent reporting it and you reading it.
How autonomous is "autonomous pentesting"?
OWASP has a draft answer. The Autonomous Penetration Testing Standard (APTS) is an OWASP incubator project, at version 0.1.0 when we checked on 26 September 2026. It covers eight domains, from scope enforcement and safety controls to human oversight, manipulation resistance and reporting, with three compliance tiers (Foundation, Verified, Comprehensive). It also defines four autonomy levels:
Level
What the human does
What the system may do
L1 Assisted
Commands every action
Runs one technique per command, no chaining
L2 Supervised
Approves at every phase boundary
Chains techniques within one phase, proposes next steps
Manages multi-target campaigns with dynamic scope and adaptive strategy
Most tools marketed as agentic for web apps and APIs sit somewhere in the middle: the agents run whole attack chains on their own, but only inside a scope a person approved. When a vendor says "autonomous", ask which APTS level they mean and where a person checks the output.
Where does the human sit?
Before the run, a person verifies they own the target and approves the scope. No traffic goes out before that. During the run, platform code (not the model) enforces the scope, the rate limits and the budget. After the run, the report is checked.
At Barrion, a security engineer reviews every report from Standard level up, with review time growing from 1 hour at Standard to 8 hours at Maximum. Light runs have no expert review, which makes them fast and cheap enough for frequent checks between reviewed runs.
What are the limits of agentic pentesting?
GUI-heavy flows. Drag-and-drop editors, canvas apps and long multi-screen wizards are harder for agents than plain HTTP requests.
New business logic. An agent can test whether a discount applies twice. It's less likely to spot that your refund policy, as coded, lets a reseller profit from a rule nobody wrote down.
Scope. Agentic web testing isn't an internal network test, an Active Directory assessment or social engineering.
That's why we don't say agents replace pentesters. See AI vs manual pentesting for the longer answer.
How does Barrion do it?
A root agent coordinates one specialist agent per testing area, working in parallel waves inside an isolated sandbox built for each engagement. Deeper levels run more agents at once: 3 at Light, 5 at Standard and 20 at Deep, up to 100 at Maximum. For API targets, the order changes so API discovery, access control and authentication go first, and browser-only areas are left out.
Findings then pass three checks. The evidence request is replayed against your app, and anything that doesn't reproduce is dropped. A finding with no way to replay it stays in the report as a lower-confidence lead, not a confirmed finding. For findings worth proving, a separate proof-of-concept agent runs its own test against the live target and returns confirmed, refuted or inconclusive. An AI review then removes false positives and corrects overstated severities. Tests are mapped to all 97 OWASP WSTG v4.2 cases.
The same run can repeat on a schedule, with each finding labelled new, still open, resolved or regressed. That's continuous AI pentesting. For how fast a run gets to its first proven finding, see Time to Proof. Product facts are on the facts page.
Agentic pentesting is a penetration test run by AI agents that choose their own next action. Instead of working through a fixed list of checks, the agents read how the target responds, adjust their approach and chain requests together, much as a human tester would. A person still sets the scope and, in a well-run service, reviews what the agents report.
Is agentic pentesting the same as autonomous pentesting?
Mostly, yes. Autonomous pentesting usually means the same thing: agents test without a person approving each step. OWASP's Autonomous Penetration Testing Standard splits autonomy into four levels, from L1 Assisted, where an operator commands every action, to L4 Autonomous, where a human reviews periodically. Autonomous testing doesn't mean nobody is involved. At Barrion, a security engineer reviews every report from Standard level up.
How is agentic pentesting different from an AI-assisted scanner?
An AI-assisted scanner runs the same checks it always did and uses a language model to summarise or rank the output. An agentic system lets the model decide what to test next based on what it just saw. The difference shows up in multi-step issues, like reaching another user's data through a second account, which a fixed check list rarely finds.
Can agentic pentesting replace human pentesters?
Not entirely. Agents cover a wide surface quickly and repeat the work as often as you like. People are still better at GUI-heavy flows and at business logic nobody has seen before. Most teams run agents continuously and bring in people periodically for the judgement-heavy paths.
Are agentic pentests safe to run against production?
They can be, if the platform enforces limits outside the model. Barrion's runs are rate-limited and non-destructive, scope is approved before any traffic is sent, and staging targets are supported. OWASP's draft APTS puts scope enforcement and a working stop control in its Foundation tier, so ask any vendor how it handles both.
How do you know an AI agent's finding is real?
You ask for proof. At Barrion, a finding's evidence is replayed against the live app before the report, and findings that don't reproduce are dropped. A finding that can't be replayed is kept as a lower-confidence lead, never marked confirmed. For findings worth proving, a separate agent runs a proof of concept and returns confirmed, refuted or inconclusive. From Standard up, a security engineer reviews the report too.
See an agentic pentest on your own app.
Book a call to scope continuous testing, or start a single run today and read the report.