What LLMs Can Do in Pentesting and Code Security
In this blog post I will discuss the use of large language models (LLMs) in penetration testing.
This is based on my prior research in the field with AutoPentest and the paper I wrote on it. It is also based on my independent industry observations. Although I do not actively work as a penetration tester, the potential of applying LLMs to this field has been fascinating to me since my initial research in June 2024.
I will start by outlining the high-level process of a pentest and discuss the stages in which LLMs can be applied. We will then look at the question of autonomy and human in the loop.
Next we dive into what kind of architecture choices can be made for your agents and their harness. I will give the AutoPentest architecture as an example that includes some of these choices.
We will then see how such a setup can be measured with existing benchmarks and finally conclude.
This post expands on my German OWASP Day talk "What LLMs Can Do in Pentesting and Code Security". There is also a talk recording available at media.ccc.de and in May 2026 I published a blog post evaluating OpenAI Codex Security in Practice
Stages of Penetration Testing
Below are the high-level stages of a pentest. During pre-engagement the scope of the test, duration, objectives and
legal stuff is
sorted out. During information gathering the target is enumerated, including scanning for vulnerabilities. Next
vulnerabilities and likely areas of successful exploitation are prioritized and further analyzed. During exploitation
real attacks are simulated to verify previously discovered vulnerabilities. Finally, a professional report is created
that includes findings and remediation recommendations.

LLMs can assist only limited in the pre-engagement, as this part usually relies heavily on human discussions and alignment. Reporting capabilities can be supported by LLMs although larger parts of report structure and format can be implemented with deterministic automation. The most interesting stages lie in the middle, where LLMs can assist in finding and exploiting vulnerabilities.
Pentest Type
A pentest can be roughly categorized into one of three categories (black box, grey box, white box) based on the assumptions under which the test is performed and which context is shared with the tester.
| Pentest type | Typical tester perspective | What context the LLM has to work with |
|---|---|---|
| Black box | External attacker with no privileged knowledge | Only public target information, observed behavior, and anything discovered during reconnaissance. |
| Grey box | Authenticated user, partner, insider, or attacker after initial foothold | Partial internal context such as low-privilege or role-based credentials, API documentation, selected architecture diagrams, and sometimes scoped internal endpoints. |
| White box | Internal assessment with full design visibility | Full internal context including source code, architecture and data-flow documentation, configuration details, credentials, schemas, build/deployment context, and sometimes logs or test data. |
System Autonomy and Human in The Loop
Below are the autonomy levels that I think make sense to reason about the behavior of both LLM/agent and the human during penetration tests. LLMs can serve anywhere from a mere assistance role to running with full autonomy only reporting consolidated findings back to the human user.
| # | Autonomy Level | LLM/Agent Behavior | Human Role | Human Review |
|---|---|---|---|---|
| 0 | Manual LLM assistance | LLM only provides suggestions in an external chat; no direct tool or shell integration. | Operator | Review every response, copy commands and execute manually. |
| 1 | Assisted command generation | LLM generates command(s) that can be run in a shell with less friction, but the human still decides what to run. | Operator with lightweight automation | Review every response and execute commands manually. |
| 2 | Subtask autonomy | Agent autonomously performs bounded subtasks such as enumeration, parsing, or exploit preparation under a user-provided plan. | Collaborator | Human steers with a high-level plan and reviews subtask results. |
| 3 | Goal-directed autonomy with safety gates | Agent follows a high-level objective across multiple steps and can continue execution unless a gated action requires approval. | Approver | Human reviews only selected safety-critical or irreversible actions. |
| 4 | Full autonomy | Agent pursues the high-level goal end-to-end without interruption and reports outcomes afterward. | Observer | Human reviews only final findings; no safety checkpoint during execution. |
The more autonomy a LLM system is given, the more security risks and financial risks are present and need to be addressed with proper safeguards.
Architecture Options
When building or choosing an LLM setup for pentesting, you can think in three layers: the harness and agentic loop, the knowledge sources it can use, and the safeguards that keep it safe to run.
Harness and agentic loop
The harness is the “engine” that runs the LLM, gives it tools, and controls how it acts.
- Purpose-built for pentesting – Systems designed specifically for offensive security, such as CAI or AutoPentest. They often ship with pentest-oriented tools, workflows, and safety gates.
- Agent frameworks or custom builds – You can build on general agent frameworks like LangChain, CrewAI, or AutoGen, or write a fully custom harness. Frameworks speed up development; custom code gives you maximum control.
- General-purpose coding agents – Tools like Claude Code, Codex CLI, Copilot CLI, or Gemini CLI. These are not pentest-specific but can be adapted with the right tools and constraints.
- Custom skills, tools, and MCP servers – Add pentest-specific capabilities (recon tools, exploit runners, parsers) via custom tools or Model Context Protocol (MCP) servers.
- Model choices – Use a single model or combine several (for example, one for reasoning, one for code, one for short tasks). You can pick specialized “cyber” models or general-purpose models.
- Orchestration – Run a single agent or multiple agents that collaborate (e.g., one for reconnaissance, one for exploitation, one for reporting).
Knowledge sources
These are the inputs the LLM can draw on while working.
- Training data and fine-tuning – What the model already knows from pretraining, plus any domain-specific fine-tuning you add (e.g., on pentest reports or exploit write-ups).
- User input during execution – Prompts, plans, constraints, and corrections you provide while the agent runs.
- Tools – External capabilities like web search, DNS lookups, or vulnerability databases that the agent can call.
- Retrieval Augmented Generation (RAG) – The agent retrieves relevant document chunks (e.g., internal architecture docs, past reports, exploit notes) and injects them into context at generation time.
- LLM wiki – The agent incrementally builds and maintains a structured knowledge base (a “wiki”) of findings, patterns, and target details, and consults it over time.
Essential safeguards
Strong containment and oversight is needed over your system.
- Attacker environment virtualization – Run the agent in isolated VMs or containers so mistakes or malicious payloads do not affect your host or production systems.
- Network traffic firewall – Restrict outbound and lateral traffic from the agent environment to only what is required for the engagement.
- Tool and command permissions – Define which tools and commands the agent may use, and require human approval for high-risk actions (e.g., exploit execution, data exfiltration simulations).
- Budget limits – Set hard or soft limits on runtime, token usage, number of tool calls, or cost to avoid runaway agents.
- Tracing and observability – Log prompts, tool calls, decisions, and outputs so you can audit behavior, debug failures, and improve the system over time.
AutoPentest Architecture (Example)
AutoPentest implements a multi-agent, human-in-the-loop architecture for web-focused penetration testing. The human user
has a minimal but safety-critical role, while several specialized LLM agents can be picked to perform the offensive
workload under a central planner.

High-level flow
- User – Starts a run by providing a target host (IP or domain) via CLI. By default, the user is prompted to allow or deny each shell command; otherwise they mainly observe progress and receive final findings.
- Service discovery – Before any LLM reasoning, AutoPentest runs
nmapto enumerate ports, services, and OS, then queries the NIST NVD API for known CVEs on identified CPEs. These results are stored and passed as context to all agents.
Agent roles
AutoPentest uses three layers of LLM agents with different system prompts, orchestrated under GPT‑4o:
- Planner – Creates and maintains a high-level, multi-step pentest plan. After each step is executed or aborted, it re-plans based on new observations, adapting or discarding infeasible ideas.
- Supervisor – Given the next planned step and current context, the Supervisor decides which Specialised Worker should execute that step.
- Specialised Workers – Execute individual plan steps with focused expertise aligned to the OWASP Top 10 (2021) and privilege escalation. Workers can use tools autonomously, but shell commands are subject to human review if enabled.
Deterministic components and safety
Service discovery (nmap + NVD), tool implementations, and RAG retrieval are deterministically implemented (hard-coded
logic with repeatable outputs). These modules provide stable capabilities and context, while the LLM agents focus on
reasoning and decision-making.
Human-in-the-loop is kept simple but effective: the user defines the target and, by default, explicitly approves or denies each shell command, providing a safety gate around high-risk actions while allowing multi-step autonomous offensive workflows.
Evaluation
Following timeline provides a chronological overview of how evaluation has evolved in this research field. The table below examines the currently most useful benchmarks in detail for different use cases.
I find the following benchmarks currently to be most useful in evaluating LLM performance in pentesting. Each serves a specific use case. While CyberGym, CyberGym-E2E and ExploitGym can be useful for evaluating model/harness performance on white-box testing scenarios with known source code, TermiBench and AgentCyberRange can be useful for evaluating black-box performance in larger cyber ranges. Wiz Cyber Model Arena can be used for broad model/harness comparisons with optional filtering by data set category.
| Benchmark | First public release | Best used for | Scope |
|---|---|---|---|
| TermiBench (dataset) | 2025-09-11 | Black-box IP-to-shell pentesting | 510 multi-service targets built around 30 CVEs across 25 services; useful for realistic service discovery, exploitation, and shell-oriented compromise. |
| AgentCyberRange (code) | 2026-06-12 | Realistic web-to-internal-network attacks | 110 vulnerabilities across 15 real web apps and 8 enterprise-like ranges with 156 internal hosts; useful for evaluating footholds, pivoting, and post-exploitation in realistic multi-host settings. |
| CyberGym (site) | 2025-06-03 | Source-assisted vulnerability verification | 1,507 historical vulnerabilities from 188 OSS-Fuzz projects; useful for testing whether agents can reproduce vulnerabilities from code and descriptions by generating a working PoC. |
| CyberGym-E2E (site) | 2026-06-03 | End-to-end vuln discovery, PoC, and remediation | 920 real vulnerabilities across 139 open-source projects; useful for testing the full workflow from finding a bug to generating a PoC and validating a fix. |
| ExploitGym (site) | 2026-05-11 | Low-level source-assisted exploit generation | 502 userspace, 181 V8 engine, and 186 Linux-kernel exploitation tasks; useful when the question is whether an agent can turn a known vulnerability into a working exploit. |
| Wiz Cyber Model Arena (site) | 2026-02-12 | Broad model and harness comparison | 221 code, 54 API, 29 web CTF, and 23 cloud-security challenges; useful for comparing agent stacks on success, time, cost, and steps across multiple offensive domains. Only results are public, but no complete data set. |
One of the main challenges that I see with standardized public benchmarks is that they either:
- Release their data set publicly to make the benchmark verifiable. Successive frontier model training on that data set is likely. This can result in overfitting and great performance on such datasets, while lacking generalized capability on untrained real world target environments or vulnerabilities.
- Keep their data set private (e.g., Wiz Cyber Model Arena). This makes benchmark claims unverifiable by the public.
I therefore argue that it can be useful to create your own private standardized benchmark data set that is reasonably close to the targets that you want to do pentests against longterm. Run any paid vendor or open source products against it before choosing one for your use case.
Conclusion
LLMs are becoming useful components of penetration-testing and code-security workflows, particularly for reconnaissance, vulnerability analysis, exploit development, and repetitive multi-step tasks. Their practical value, however, depends less on the model alone than on the surrounding harness: tool integration, deterministic components, contextual knowledge, isolation, observability, and human approval for high-risk actions.
Near-term progress is likely to come from systems that combine bounded autonomy with strong safety gates rather than from fully autonomous agents. Benchmark results should therefore be interpreted cautiously: public datasets risk training contamination and saturation, while private benchmarks limit reproducibility. For meaningful evaluation, organizations should supplement public benchmarks with private, representative test environments and measure more than success rate, including reliability, time, cost, false positives and safety violations.
The most credible path forward is not to replace experienced security professionals, but to give them systems that can explore more hypotheses, execute more routine work, and produce better-supported findings without weakening operational control.