Human-in-the-loop in autonomous AI security testing
Short definition
Human-in-the-loop (HITL) is the design principle that an autonomous AI system acts within limits set by people and hands specific decisions — here, any action that could affect a production system — back to a human for approval before it happens.
Why this matters now
Autonomy is what makes an AI red team useful; control is what makes it deployable. Boards, regulators and the EU AI Act (human oversight, Art. 14) all ask the same question: who can stop it, and who approved what it did? A well-designed HITL model answers both, with evidence.
Key points
- ▸Graduated autonomy: from passive observation only, to fully autonomous testing within a defined scope.
- ▸Approval gates: actions with potential impact on a target wait for an explicit human decision.
- ▸Stop control: operators can pause or terminate a campaign at any time.
- ▸Accountability: every approval and every action is recorded with who, what and when.
- ▸Final judgement stays human: automated verdicts are proposals an operator confirms.
How graduated autonomy works in practice
In Zero Hunt, each campaign runs at one of five autonomy levels, from passive-only observation up to full assessment reserved for scheduled maintenance windows. Each level defines what the agents may do on their own and what must wait for a person: at the lowest level any active scan needs approval; intermediate levels require approval for exploitation, credential and availability tests; only the highest level, chosen deliberately, runs without approval gates inside its scope.
When the engine proposes to run a proof of concept that the campaign level does not allow, it does not run it: it places the script in a review tab where an operator can inspect it and decide. Execution there requires a recorded consent, a fresh scope check at the moment of execution, a per-script permission that does not change the campaign level, a lock that prevents two actors touching the same host at once, and a signed record of the operator's identity.
HITL, the EU AI Act and your auditors
The EU AI Act requires high-risk AI systems to be designed so that people can oversee them effectively, understand their output and intervene or interrupt them (Art. 14). NIS2 and DORA ask for evidence that testing is controlled and proportionate.
A HITL design turns those requirements into artefacts: the autonomy level of every campaign, each approval with its author and time, each action with its evidence, and a stop control that works. For more on the operating model, see autonomous AI red team and the on-premise AI red team overview.
Goes deeper
Want this against your environment?
Book a 30-minute scoping call — we will map this directly to your current compliance scope and threat profile.