Continuous Automated Red Teaming (CART): What It Is and What It Isn't
Short definition
Continuous automated red teaming (CART) is software that runs offensive tests against your own environment on a recurring basis and after changes, attempting real attacks and reporting which ones succeed, so exposure is measured continuously rather than once a year.
Why this matters now
Vendors use CART alongside automated pentesting, BAS and AEV, often for similar products, and no standard defines it. Buyers who take the phrase literally may expect it to replace a red team exercise or a regulator-mandated test such as TLPT or CBEST, which it does not. Knowing what continuous should mean in practice, and what safety controls unattended attacks need, separates a useful program from a noisy scheduler.
Key points
- ▸CART is a market term with no standard or regulatory definition; NIST's glossary defines a red team as a group of people, not a tool.
- ▸In practice CART means automated or autonomous penetration testing run continuously, sometimes with attack surface discovery in front of it.
- ▸A red team exercise is objective-driven, human-led and broader, including social engineering and physical access; CART covers part of it.
- ▸BAS validates controls against catalogued techniques, automated pentesting proves attack paths, and AEV is Gartner's category that covers both; CART usually sits on the pentesting side.
- ▸Continuous should mean scheduled runs, runs triggered by change, a retest after every fix and a record of what changed since the last run.
- ▸CART does not replace TLPT, CBEST or other tests where a regulator names the tester; it gives you evidence between them.
What CART means, and why the term is loose
Continuous automated red teaming, or CART, describes software that attacks your own environment on a recurring basis, the way an adversary would, and reports what succeeded. The three words each carry a promise: continuous (not a yearly snapshot), automated (no person driving each step) and red teaming (adversary emulation rather than a checklist).
There is no standard or regulatory definition of CART. The NIST CSRC glossary has no entry for it, and its definition of a red team, from CNSSI 4009-2022, describes people, not software: “a group of people authorized and organized to emulate a potential adversary's attack or exploitation capabilities against an enterprise's security posture”, whose objective is to improve security “by demonstrating the impacts of successful attacks and by demonstrating what works for the defenders (i.e., the Blue Team) in an operational environment”.
So CART is a vendor and market term, and products sold under it vary. Most are automated or autonomous penetration testing run on a schedule, often combined with discovery of the external attack surface. Judge a product by what it does, not by the label.
CART vs red team exercises, BAS, automated pentesting and AEV
- Red team exercise. An objective-driven campaign run by people, which can include social engineering and physical access and is often unannounced to the defenders. It tests the organization, its people and its detection and response, over weeks. CART automates part of the technical attack, not the whole exercise.
- Breach and attack simulation (BAS). Replays catalogued, controlled attack scenarios, often through agents you install, to check whether your controls block and detect known techniques. It answers a defense question and does not look for paths nobody wrote a scenario for. See BAS vs automated pentesting vs AEV.
- Automated penetration testing. Discovers assets, attempts real exploitation, reuses credentials, moves laterally and reports the paths that worked, with evidence. Autonomous tools let AI agents choose the next step. This is what most CART products do.
- Adversarial Exposure Validation (AEV). Not a technique but Gartner's market category for technologies that deliver “consistent, continuous and automated evidence of the feasibility of an attack”; Gartner states it replaces BAS and automated penetration testing and red teaming technology from its 2023 Hype Cycle. CART products fall within it. See Adversarial Exposure Validation (AEV).
- Continuous threat exposure management (CTEM). A program, not a product. CART supplies evidence for its validation stage. See CTEM.
One more confusion: “AI red teaming” often means testing AI models, not networks. A CART product attacks your infrastructure; a model red-teaming tool attacks your language model.
What “continuous” should mean in practice
A product that can run on a schedule is not yet a continuous program. Ask for, and plan, these five things.
- A cadence matched to the rate of change. Internet-facing systems and identity infrastructure change often and are attacked first; test them more often than a stable internal segment. Rules set some floors: DORA's ICT risk management RTS requires automated vulnerability scanning of assets supporting critical or important functions at least weekly, and PCI DSS requires scans every three months and a penetration test every 12 months.
- Runs triggered by change. A new internet-facing service, a firewall change, a major release, a merger or a newly exploited vulnerability in software you run should start a test without waiting for the next slot. PCI DSS 11.3 and 11.4 and NYDFS 500.5 already require tests or scans after significant or material changes, and the ACN measures require a test before a relevant system goes live. For newly exploited vulnerabilities, see the KEV-driven emergency patch window.
- A retest after every fix. A finding closes when the same attack fails, not when a ticket is marked resolved. PCI DSS 11.4.4 requires penetration testing to be repeated to verify corrections.
- A record of what changed. New findings since the last run, findings that came back after being fixed, and assets not tested in the period. Without it, a continuous tool produces the same report every week.
- Someone who acts on it. Continuous testing pays off only if findings reach an owner with a deadline. Plan the remediation capacity before the cadence.
If the product also has to answer whether your SOC saw the attack, correlate its timeline with your alerts; automated pentesting tells you what worked, not whether it was detected.
Safety and human approval when nobody is watching
A consultant running a red team exercise is watching every step. A continuous tool often runs at night, on production, with no one in front of it, and an autonomous one decides its own next action. The safety controls have to be in the engine:
- Scope checked on every action, including addresses inside code the tool generates, at the moment of execution. The tool's own host excluded. Newly discovered hosts added to scope only by explicit opt-in, and logged.
- Approval gates by class of action. Decide which actions wait for a person: exploitation, credential attacks, anything that changes the state of a target, denial-of-service tests. For unattended runs, choose deliberately what the tool may do on its own and what should wait until someone is available.
- An emergency stop that holds. One action that halts everything, survives a restart and is cleared only by an operator.
- Rate limits and testing windows, so scheduled runs do not saturate fragile systems or fall inside business-critical hours.
- A tamper-evident record of every action by agents and operators, so an incident during a run can be reconstructed.
For the reasoning behind approval gates, see human in the loop in AI security testing.
What CART does not replace
- Regulator-mandated threat-led tests. DORA requires the financial entities identified by their competent authority to carry out threat-led penetration testing at least every three years, with an active red team phase of at least 12 weeks and strict rules on testers: the threat intelligence provider is always external, internal testers need the authority's approval and external testers are required every three tests. See TLPT and the DORA TLPT engagement playbook.
- CBEST. The Bank of England, PRA and FCA intelligence-led assessment is run by CBEST-accredited threat intelligence and penetration testing providers when the regulators request it or agree to it, and includes about 14 weeks of penetration testing. See UK CBEST and CAF.
- Tests that name a qualified or accredited party, such as PCI ASV scans or the annual NYDFS penetration test by a qualified party. See penetration testing requirements by regulation.
- Social engineering and physical access, which are part of real intrusions and of red team exercises and lie outside what a network testing tool does.
- Business-logic testing and judgment. Abusing an approval workflow or a pricing rule, deciding what may be tested and what a finding means for the business, remain work for people.
What CART changes is the time between those exercises: instead of arriving at a TLPT or a CBEST with a year-old picture, you arrive with continuous evidence of what was tested and what was fixed. That makes the mandated test a confirmation rather than a discovery; it does not make it optional.
Questions to ask a CART vendor
- What do you actually execute? Real exploitation attempts, simulated techniques, or modeled paths? For which findings do you have proof?
- How do you decide the next step? A fixed sequence or agents that adapt? What stops an agent from leaving scope?
- What triggers a run? Schedules only, or also changes, new assets and new vulnerabilities? Can a single finding be retested on demand?
- What happens on unattended runs? Which actions wait for approval, who is notified, and what happens to a pending approval when nobody answers?
- Can I stop it? Show the emergency stop working, including after a restart.
- Black box, gray box or both? How are the credentials I provide stored and used?
- Where does it run and what leaves my network? SaaS, cloud control plane with a local agent, or fully on-premise? Where does the AI model run? Can it run air-gapped?
- What is the evidence? Can I prove records were not altered? Can I export them for an auditor, mapped to the frameworks I report against?
- What does continuous cost? Under per-test or usage-based pricing, every retest adds cost; model a year of weekly runs, not one test. See automated pentesting cost.
- What do you claim not to replace? A vendor that says its product replaces TLPT, CBEST or a qualified tester is overselling.
Zero Hunt's approach (vendor section)
The sections above apply whatever product you choose. This one describes how Zero Hunt, the product behind this site, answers them.
- What it executes. Zero Hunt is an autonomous AI red team for networks and infrastructure: a controller coordinates ten specialized agents that attempt real exploitation and prove which exposures are exploitable, with evidence per finding. It is not a BAS tool and does not replay email or malware delivery scenarios against gateways.
- Continuous. Campaigns run once, daily, weekly, monthly or on a custom schedule, and an operator can start one after a change. A campaign re-run after a fix shows closure with signed evidence.
- Human in the loop. Five autonomy levels decide which actions wait for an operator, from any active scan at the lowest level to no approval gates at the highest, which is chosen deliberately. Exploit verification is enabled only at the two highest levels; below them an operator can run a single proof of concept after a written, named consent, with scope re-checked at that moment.
- Safety. Scope is validated on tool calls and in generated scripts, the appliance's own addresses are excluded, runtime scope expansion is opt-in and logged, and the emergency stop persists across restarts until an operator clears it.
- Evidence. Every attack attempt is an Ed25519-signed record in a SHA-256 hash chain per campaign; exported reports and bundles are ECDSA-signed; findings map to 34 compliance frameworks worldwide.
- On-premise, private AI. Its own models run on the appliance, no customer data leaves it, and air-gapped mode disables public OSINT and public source-code downloads.
Zero Hunt is not named by Gartner as an AEV vendor, and it does not replace TLPT, CBEST or the qualified testers other rules require. Read about the on-premise AI red team, or book a 30-minute readiness call to map continuous testing to the rules you answer to.
Sources
- Red team, glossary definition (NIST CSRC, from CNSSI 4009-2022)
- Penetration testing, glossary definitions (NIST CSRC)
- Adversarial Exposure Validation, market definition (Gartner Peer Insights)
- Regulation (EU) 2022/2554 (DORA), Articles 24 to 27 (EUR-Lex)
- Commission Delegated Regulation (EU) 2025/1190, RTS on TLPT (EUR-Lex)
- Commission Delegated Regulation (EU) 2024/1774, RTS on the ICT risk-management framework (EUR-Lex)
- CBEST Implementation Guide 2024 (Bank of England)
- PCI Data Security Standard (PCI SSC)
The regulatory summaries repeat the detailed guides linked in the text, which cite their full primary sources.
Goes deeper
- Automated penetration testing: how it works and its limits →
- BAS vs automated pentesting vs AEV →
- Adversarial Exposure Validation (AEV) →
- TLPT: threat-led penetration testing →
- UK CBEST and CAF penetration testing →
- Human in the loop in AI security testing →
- Vulnerability assessment vs penetration testing (VA/PT) →
- On-premise AI red team on private AI →
- Book a 30-minute readiness call →
Want this against your environment?
Book a 30-minute scoping call — we will map this directly to your current compliance scope and threat profile.