← Learn
Definition12 min read

Automated Penetration Testing: How It Works, Limits, and How to Evaluate Tools

Short definition

Automated penetration testing is software that carries out the steps of a penetration test (discovery, exploitation attempts, privilege escalation, lateral movement, reporting) without a person driving each step, and shows which weaknesses an attacker could actually use.

Why this matters now

Most organizations still test once or twice a year while their networks change every week, and attackers automate too: in March 2025 Gartner predicted that by 2027 AI agents will cut the time it takes to exploit account exposures by 50%. Automation closes the gap between tests. It does not remove the need for qualified people, and several regulations still say so explicitly.

Key points

  • ▸Manual testing is driven by a person, automated testing runs predefined attack sequences, and autonomous (agentic) testing lets AI agents choose the next step from what they find.
  • ▸The output that matters is proof: evidence that a weakness was exploited in your environment, not a list of possible vulnerabilities.
  • ▸Black-box tests start from what an outsider sees; gray-box tests start with credentials or source code, like an insider or a compromised account.
  • ▸Safety lives in the engine: scope checked on every action, approval gates for intrusive steps and an emergency stop that holds.
  • ▸Automation does not replace people for business-logic flaws, social engineering, scoping decisions or tests a regulator requires from qualified testers, such as TLPT.
  • ▸Evaluate tools on deployment model, data egress, exploit proof, human approval, evidence, compliance mapping and pricing model.

Manual, automated and autonomous pentesting

NIST defines penetration testing as “a test methodology in which assessors, typically working under specific constraints, attempt to circumvent or defeat the security features of a system”. The key word is attempt: a vulnerability scan that compares software versions against a catalog is not a penetration test, however automated it is. What changes between the three approaches is who decides the next attempt.

  • Manual penetration testing. A tester, or a team, plans and runs every step. It brings judgment and creativity, and it is the only option for some targets, but it is expensive, happens a few weeks a year and depends on the skill of the individuals.
  • Automated penetration testing. Software runs predefined attack sequences: discover hosts, match services to known weaknesses, try known exploits and credential attacks, and follow the paths that open up. It is repeatable and can run on a schedule, but it is limited to the techniques its authors encoded.
  • Autonomous (agentic) penetration testing. AI agents read what they find, form a hypothesis, choose or write the next action and keep going until the hypothesis is proven or disproven. They can handle situations no playbook anticipated, which also makes their behavior harder to predict. The safety controls described below matter most here.

The market uses these labels loosely. Look at the mechanism instead: does the tool follow a fixed sequence or decide its next step, and does it actually attempt exploitation or only infer it?

How an automated penetration test runs

Most tools, whatever they call themselves, follow the same cycle:

  1. Scope and rules of engagement. Address ranges, domains and applications in scope; what is excluded; testing windows; which actions need approval; who can stop the test.
  2. Discovery. Hosts, open services, software versions, web applications, identities and trust relationships.
  3. Analysis. Matching what was found with known weaknesses, misconfigurations and default or weak credentials, and deciding what to try first.
  4. Exploitation attempts. Trying to use a weakness for real, within the limits the operator set. This is the step that separates a penetration test from a scan.
  5. Post-exploitation. From a foothold: privilege escalation, credential harvesting, lateral movement, and the chain that leads to the systems that matter.
  6. Evidence and reporting. Each finding with the commands, responses and artifacts that prove it, a severity and remediation guidance.
  7. Retest. After a fix, the same proof runs again to show that the path is closed.

Because steps 2 to 7 are software, they can repeat every week or after every significant change instead of once a year. That is the main value of automation: frequency and consistency, not a better single test.

Black box and gray box

The starting knowledge decides which attacker the test represents. A black-box test starts with nothing but the scope, the view of an outside attacker. A gray-box test starts with partial knowledge: the credentials of a standard user, documentation, or the source code of the software in use. That is the view of an insider, or of an attacker who has already phished an account or done their homework.

The two answer different questions. Black box measures what is exposed; gray box measures what an informed attacker can do behind the login. Many tools also offer an assumed-breach start from a host inside the network. Ask which modes a tool supports and how the credentials you provide are stored and used. See black-box vs gray-box penetration testing.

The safety controls to require

An automated tool acts on production systems, and an agentic one writes its own actions. Safety cannot rest on a scope document; the engine has to enforce it. Require at least:

  • Scope validation on every action. The target of each tool call, and any address inside code the tool generates, checked against the authorized scope at the moment of execution, not only when the campaign starts. The tool's own host excluded. Scope extended to newly discovered hosts only by explicit opt-in, and logged.
  • Approval gates. Graduated autonomy, where you decide which classes of action wait for a person: exploitation, credential attacks, anything that can change the state of a target, denial-of-service tests. Full autonomy should be a deliberate setting, not the default.
  • An emergency stop. One action that halts all operations, survives a restart of the tool and can only be cleared by an explicit operator decision.
  • Rate limits and testing windows, so discovery and exploitation do not saturate fragile systems or run outside agreed hours.
  • Isolation of generated code. Exploit code runs in a contained environment that cannot reach the tester's own infrastructure.
  • A tamper-evident record of every action, by agents and by operators, so you can reconstruct what happened if something breaks.

Ask vendors to show these controls working rather than describe them. A simple test during a pilot: start a campaign, trigger the emergency stop, restart the service and check that it stays stopped.

What automation still cannot replace

Automation raises frequency and consistency. It does not make human testers obsolete, and a tool that claims otherwise is overselling.

  • Business-logic flaws. Abusing an approval workflow, a refund process or a pricing rule requires understanding what the application is supposed to do. Tools are improving here, but people still find most of these flaws.
  • Social engineering and physical access. Phishing, help-desk pretexting and walking into a building are part of real intrusions and of threat-led tests, and they are outside what a network testing tool does.
  • Scoping and risk decisions. Which systems may be tested, how much risk to production is acceptable and whether a finding matters to the business are decisions for accountable people.
  • Interpretation. Turning findings into a narrative for the board, a regulator or an auditor, and deciding what to fix first against business priorities.
  • Absence of evidence. A clean automated run shows that the techniques the tool tried did not work. It does not prove that nothing else would.

The model that works is a combination: automated or autonomous testing running continuously, qualified people for the periodic deep tests and for what tools cannot reach, and the tool's findings as input to those people rather than a substitute for them.

When regulators still require human testers

Several rules set requirements on who performs the test, and an automated tool does not meet them on its own.

  • DORA threat-led penetration testing (TLPT). Financial entities identified by their competent authority must carry out TLPT at least every three years, on live production systems supporting critical or important functions (Regulation (EU) 2022/2554, Art. 26). Testers must be of the highest suitability and reputability; possess technical and organisational capabilities and specific expertise in threat intelligence, penetration testing and red team testing; be certified by an accreditation body in a Member State or adhere to formal codes of conduct or ethical frameworks; provide independent assurance or an audit report on the management of the test's risks; and be covered by professional indemnity insurance (Art. 27(1)). Using internal testers requires the authority's approval, external testers must be contracted every three tests, and significant credit institutions may use only external testers (Art. 26(8) and 27(2)). See TLPT and the DORA TLPT engagement playbook.
  • NYDFS 23 NYCRR 500.5(a)(1). Penetration testing from both inside and outside the information systems' boundaries “by a qualified internal or external party at least annually”. See the NYDFS Part 500 playbook.
  • CMMC Level 3, CA.L3-3.12.1e. Penetration testing at least annually or when significant security changes are made to the system, “leveraging automated scanning tools and ad hoc tests using subject matter experts”. The requirement names both. See the CMMC Phase 2 playbook.

The reading is the same in each case: automation gives you evidence between the mandated exercises and helps you arrive at them prepared, while the mandated exercise stays with qualified people.

How to evaluate automated pentesting tools: a checklist

Use these questions in a request for proposal or a pilot. They separate tools that look alike in a demo.

  1. Deployment model. Is it SaaS, a cloud control plane with an agent in your network, or fully on-premise? Can it run with no internet connection at all? Gartner notes that products in its Adversarial Exposure Validation category are generally delivered as SaaS, with or without on-premises agents, so do not assume on-premise.
  2. Data egress. What leaves your network: network maps, findings, recovered credentials, screenshots, prompts to an AI model? Where does the AI model run, and who else processes that data? Is any of it used to train models?
  3. Exploit proof. Does the tool actually attempt exploitation, or infer exploitability from versions? What does the evidence for a finding contain? How are severities set for findings it could not prove?
  4. Human approval. Which actions wait for a person at each autonomy level? Can you set it per campaign? Is every approval recorded with a name and a time?
  5. Safety controls. Scope checked per action, emergency stop, rate limits, testing windows, and no denial-of-service tests unless you enable them.
  6. Evidence. Can you prove the test records were not altered? Are they signed, chained and verifiable offline? Can you export them for an auditor?
  7. Compliance mapping. Which frameworks does it map findings to, and how? A tool that claims to replace a legally required test is a red flag.
  8. Coverage and modes. Internal network, Active Directory, web applications, APIs, cloud, OT; black box and gray box; how credentials are handled.
  9. Retesting. Can every fix be retested on demand with the same proof, and is that included in the price?
  10. Pricing model. Per asset or IP, per test, usage-based credits or tokens, subscription, or an appliance. Model what continuous testing will cost, not what one test costs: usage-based pricing grows with exactly the testing you want more of.
  11. A pilot on your own network. Pick a segment with weaknesses you already know about and see what the tool finds, what it proves and what it misses.

Zero Hunt's approach (vendor section)

The sections above are meant to be useful whatever tool you choose. This one describes how Zero Hunt, the product behind this site, answers the checklist.

  • Deployment and data. Zero Hunt is an on-premise appliance that runs an autonomous AI red team for networks and infrastructure on private AI: its own ZeroHunt Apex models on the appliance GPU, with no external AI service. No customer data leaves the appliance. A connected appliance fetches signed updates and public threat intelligence; air-gapped mode turns off public OSINT and public source-code downloads.
  • How it tests. A controller coordinates ten specialized agents: reconnaissance, vulnerability analysis, exploitation, web exploitation, credentials, post-exploitation, pivoting, ATT&CK tactic planning, source analysis and reporting. Campaigns run black box by default, and gray box through authenticated testing with credentials you provide or through analysis of the source of the exact software version in use.
  • Human in the loop. Five autonomy levels. At the lowest, any active scan waits for approval; at the next, exploit and credential tests; then exploit and denial-of-service tests; then denial-of-service tests only; the highest has no approval gates and is chosen deliberately. Exploit verification is enabled only at the two highest levels. Below them, an operator can run a single proof of concept after a written, named consent, with the scope re-checked at that moment; the engine proposes a verdict and the operator has the final word. See human in the loop.
  • Safety. Scope is validated on tool calls and in generated scripts, the appliance's own addresses are excluded, runtime scope expansion is opt-in and logged, and the emergency stop persists across restarts until an operator clears it.
  • Evidence and compliance. Every attack attempt is an Ed25519-signed record in a SHA-256 hash chain per campaign; exported reports and bundles are ECDSA-signed. Findings map to 34 frameworks across the EU, US, UK, Canada, APAC, the Middle East, Latin America and Africa.
  • Cost. The appliance is a flat cost, with no per-token metering however many campaigns you run.

Zero Hunt does not replace the qualified testers that TLPT, NYDFS Part 500 or CMMC Level 3 require; it gives you continuous, evidence-backed testing between those exercises. Compare it with other tools on the alternatives page, read about the on-premise AI red team, or request a demo. Related: BAS vs automated pentesting vs AEV and Adversarial Exposure Validation (AEV).

Sources

Goes deeper

Want this against your environment?

Book a 30-minute scoping call — we will map this directly to your current compliance scope and threat profile.