← Learn
Definition11 min read

Vulnerability Assessment vs Penetration Testing (VA/PT): Differences and When You Need Each

Short definition

A vulnerability assessment finds and rates known weaknesses across many systems. A penetration test tries to exploit weaknesses, chains them, and proves what an attacker can actually reach. VA/PT, or VAPT, is the common name for doing both.

Why this matters now

Many rules accept either test or name only one, and buyers often pay for a scan labeled as a pentest or for a pentest where a scan would do. In Italy, NIS2 essential entities must run at least vulnerability assessment and/or penetration testing on relevant systems under the ACN measures, and PCI DSS separates quarterly scans from annual penetration tests. Knowing which question each test answers decides the budget and the evidence you can show.

Key points

  • ▸Vulnerability assessment answers what known weaknesses exist and how severe they look; penetration testing answers which of them an attacker can actually use, and how far.
  • ▸NIST SP 800-115 notes that vulnerability scanners can have a high false positive rate, while the attack phase of a penetration test exploits the vulnerability to confirm its existence.
  • ▸VA is broad, largely automated and cheap to repeat; PT is deeper, needs expertise, carries more risk to production and costs more per test.
  • ▸Regulations differ: PCI DSS requires both (11.3 scans, 11.4 pentests); Italy's ACN ID.RA-01 accepts VA and/or PT for essential entities; NYDFS 500.5 requires an annual pentest plus scans.
  • ▸Automated and autonomous pentesting sit between the two: the frequency of a scan with actual exploitation attempts, but not the judgment of a human tester.
  • ▸A clean scan or a clean pentest is not proof of security: both only show what was tried.

Definitions: what NIST means by each

The NIST CSRC glossary, quoting CNSSI 4009-2022, defines a vulnerability assessment as a “systematic examination of an information system or product to determine the adequacy of security measures, identify security deficiencies, provide data from which to predict the effectiveness of proposed security measures, and confirm the adequacy of such measures after implementation”. In practice it is carried out mostly with vulnerability scanners, which, in the words of NIST SP 800-115, identify hosts and their attributes and match them “with information on known vulnerabilities stored in the scanners' vulnerability databases”.

The same glossary defines penetration testing as “a test methodology in which assessors, typically working under specific constraints, attempt to circumvent or defeat the security features of a system”. SP 800-115 adds that it “often involves launching real attacks on real systems and data” and that most penetration tests look “for combinations of vulnerabilities on one or more systems that can be used to gain more access than could be achieved through a single vulnerability”.

SP 800-115 puts the two in different families. Vulnerability scanning is a target identification and analysis technique. Penetration testing is a target vulnerability validation technique, whose objective “is to prove that a vulnerability exists, and to demonstrate the security exposures that occur when it is exploited”. The short version: a VA produces a list of probable weaknesses, a PT produces proof.

VA/PT or VAPT is simply the name for an engagement, or a program, that does both, usually a broad assessment followed by penetration testing of the systems where exploitability matters most.

Side by side

  • Question answered. VA: which known weaknesses exist, and how severe do they look? PT: can an attacker use them, alone or chained, to reach something that matters?
  • Scope. VA: wide, often every host and application in an inventory. PT: narrower and goal-driven, such as the internet perimeter, a critical application or the path to domain admin.
  • Depth. VA: stops at detection and a severity score. PT: attempts exploitation, then privilege escalation and lateral movement where the rules of engagement allow.
  • Output. VA: a list of findings ranked by a score, often long. PT: fewer findings, each with evidence of what was exploited and the chain that led there.
  • False positives. VA: common, because detection is inferred from versions and signatures. PT: rare for exploited findings, since the exploit is the proof.
  • Risk to production. VA: low, although aggressive scans and denial-of-service checks can still disturb fragile systems. PT: higher, because it uses real exploits against real systems; SP 800-115 notes that systems “may be damaged or otherwise rendered inoperable”.
  • Skills. VA: tooling plus someone who can interpret results. PT: SP 800-115 calls it “labor-intensive” and says it “requires great expertise to minimize the risk to targeted systems”.
  • Typical frequency. VA: continuous to quarterly. PT: annual or after major changes, unless automated.

False positives, false negatives and why the difference matters

SP 800-115 is blunt about scanners: although the process is highly automated, “vulnerability scanners can have a high false positive error rate (i.e., reporting vulnerabilities when none exist)”, and an expert should interpret the results. It also names the opposite problem. Weaknesses “rarely exist in isolation”: several low-risk findings can present a higher risk when combined, scanners cannot detect such combinations, and “a more reliable way of identifying the risk of vulnerabilities in aggregate is through penetration testing”.

This is why a VA report alone often produces two kinds of waste. Teams spend time on findings that are not real or not reachable, and they deprioritize low-scored findings that together form a path an attacker would take. A penetration test cuts both, because a finding it exploited is real and the chain it followed shows which low scores matter.

Penetration testing has its own blind spot: false negatives. A test proves what worked, within its scope, time and techniques. A clean result means the techniques tried did not succeed; it does not prove that nothing else would. The same is true of a clean scan, which only means that no signature matched.

Frequency: snapshot or continuous

Because scanning is automated, a VA can run as often as the environment changes: weekly, daily, or after every change. Several rules set floors: PCI DSS requires internal and external scans at least every three months and after significant changes, and the DORA ICT risk management RTS requires automated vulnerability scanning of assets supporting critical or important functions at least weekly.

Penetration testing has traditionally been annual. When SP 800-115 was written in 2008 it said that, because of its high cost and potential impact, penetration testing of an organization's network and systems “on an annual basis may be sufficient”. Rules still use that interval as a floor: PCI DSS 11.4 and NYDFS 500.5 require at least an annual penetration test, and both require further tests or scans after significant or material changes.

The practical problem is the gap between tests. A network that changes every week is tested for exploitability once a year, and the VA findings in between are unproven. That gap is what automated and autonomous pentesting are meant to close; see below.

What drives the cost

Neither test has a list price, but the cost drivers are well known. Ask every provider to quote the same scope and compare line by line.

  • Assets in scope. The number of live hosts, applications and APIs. For a VA it is the main driver; for a PT it matters less than depth.
  • Depth and starting point. External black box only, or gray box with user credentials, source code or an assumed-breach position inside the network. Each adds tester time.
  • Manual effort. A VA is mostly tool time plus review; a PT is billed in tester-days. The GSA schedule ceiling rates summarized in our cost guide put the median ceiling rate for a penetration tester on the US federal schedule near $158 an hour.
  • Frequency and retests. A second test after remediation, or tests after each significant change, multiply the price of a point-in-time engagement.
  • Who may test. Some rules name the party: PCI DSS external scans must come from an Approved Scanning Vendor, and threat-led tests like TLPT and CBEST have their own tester requirements.
  • Reporting and evidence. A report written for an auditor or a regulator, with methodology, tester qualifications and remediation records, takes time to produce.

For public contract values of automated pentesting platforms and a three-year cost method, see what automated pentesting costs.

Which regulations ask for which

This summary repeats our verified guides; each link quotes the text and its primary source. The full map is in penetration testing requirements by regulation.

  • NIS2 (EU). The Directive never names the penetration test. Article 21(2)(e) and (f) require vulnerability handling and policies to assess the effectiveness of measures; Implementing Regulation (EU) 2024/2690 asks the digital providers it covers for a security testing policy and, where appropriate, vulnerability scans.
  • Italy, ACN measures. Measure ID.RA-01 requires essential entities to carry out, on at least their relevant systems, activities that include “at least vulnerability assessment and/or penetration test”, periodically and in any case before go-live, documented in reports. Important entities have no explicit VA/PT requirement but need a board-approved vulnerability management plan (ID.RA-08). Details in VA/PT and NIS2.
  • PCI DSS v4.0.1. Both, as separate requirements: 11.3 internal and external vulnerability scans at least every three months (external ones by an ASV), 11.4 internal and external penetration tests at least every 12 months and after significant changes, with repeated testing to verify corrections. See PCI DSS 11.3 and 11.4.
  • DORA. Every financial entity needs a risk-based testing programme; Article 25(1) lists tests from vulnerability scans to penetration testing, systems supporting critical or important functions are tested at least yearly, and those assets are scanned at least weekly. See DORA testing beyond TLPT.
  • NYDFS 23 NYCRR 500.5. An annual penetration test from inside and outside the boundaries by a qualified party, plus automated scans and a manual review of systems the scans do not cover. See the NYDFS Part 500 playbook.

The pattern: rules that care about exploitability ask for a penetration test by name, rules that are risk-based let you choose, and in that case a scan-only program is hard to defend where the question is whether a weakness can be exploited.

Black box and gray box

Both tests can start from different levels of knowledge, and the choice changes what they measure. A black-box test starts from the outside with nothing but the scope; a gray-box test starts with partial knowledge, such as a standard user's credentials, documentation or source code.

For a VA, the gray-box equivalent is an authenticated scan, which logs in and reads installed versions and settings instead of guessing them from the network; PCI DSS 11.3.1.2 now requires authenticated internal scans. For a PT, gray box represents an insider or an attacker who already phished an account, and it usually reaches further than a black-box test in the same time. A mature program uses both. See black-box vs gray-box penetration testing.

Where automated and autonomous pentesting sit

Automated penetration testing takes the loop of a penetration test (discovery, exploitation attempts, post-exploitation, evidence, retest) and runs it in software. Autonomous, or agentic, tools let AI agents choose the next step from what they find instead of following a fixed sequence.

On the VA/PT spectrum they sit in between:

  • Like a VA, they run on a schedule, cover many assets and cost little to repeat.
  • Like a PT, they attempt exploitation and chain weaknesses, so a finding comes with proof rather than a version match, and they can retest a fix with the same proof.
  • Unlike a human PT, they are weak on business-logic flaws, do not do social engineering or physical access, and do not meet rules that name the tester, such as TLPT or CBEST.

The model that works for most programs: continuous VA for breadth, automated or autonomous pentesting for frequent proof of exploitability, and human-led penetration tests for depth, for what tools cannot reach, and wherever a rule requires a qualified or accredited tester.

A short decision checklist

  1. Do you need an inventory of known weaknesses across everything, repeatedly? That is a VA, and it should run continuously or at least at the interval your rules set.
  2. Do you need to know whether an attacker can reach specific systems, or to prioritize remediation by proven exploitability? That is a PT, human or automated.
  3. Does a rule name the test? PCI DSS needs both; NYDFS needs an annual pentest; ACN ID.RA-01 accepts VA and/or PT but expects the choice to follow your vulnerability management plan.
  4. Does a rule name the tester? ASV scans, TLPT and CBEST providers are not optional, and no tool replaces them.
  5. How often does the environment change? If it changes weekly, an annual pentest leaves most of the year unproven.
  6. Which attacker are you modeling? Outsider only (black box) or also an insider or phished user (gray box)?
  7. Can the systems take it? Fragile or OT systems may need passive assessment or a documented exception with compensating measures instead of exploitation.
  8. What evidence will you need? Reports with methodology, tester, findings, remediation and retest results, kept for the longest retention period that applies to you.

Zero Hunt's approach (vendor section)

The sections above apply whatever tool or provider you choose. This one describes how Zero Hunt, the product behind this site, fits.

  • VA and PT in one campaign. Zero Hunt is an autonomous AI red team for networks and infrastructure. Its campaigns find vulnerabilities and then try to exploit them, so each finding records whether it was proven exploitable in your environment.
  • Black box and gray box. Campaigns run black box by default, and gray box through authenticated testing with credentials you provide or through analysis of the source of the exact software version in use.
  • Continuous, with retests. Campaigns run once, daily, weekly, monthly or on a custom schedule, and can be re-run after a fix so closure is shown with signed evidence.
  • Human in the loop. Five autonomy levels decide which actions wait for an operator; exploit verification is enabled only at the two highest levels, and below them an operator can run a single proof of concept after a written, named consent.
  • Evidence. Every attack attempt is an Ed25519-signed record in a SHA-256 hash chain; findings map to 34 compliance frameworks. The engine does not map to Italian ACN measure codes, so linking its reports to ID.RA-01 stays in your documentation.
  • On-premise, private AI. It runs on an appliance with its own models; no customer data leaves it, and it can run air-gapped.

It does not replace a PCI ASV, the qualified testers some rules name, or TLPT and CBEST providers. Read about the on-premise AI red team, or book a 30-minute readiness call to check which tests your regulations require and where your evidence has gaps.

Sources

Goes deeper

Want this against your environment?

Book a 30-minute scoping call — we will map this directly to your current compliance scope and threat profile.