Abstract
Recent advances in large language models (LLMs) have opened new opportunities for automating complex tasks in Cybersecurity, including offensive operations. However, most existing approaches to LLM-assisted penetration testing rely on human input or scripted interactions. This work explores a fully autonomous, local-agent framework for end-to-end penetration testing, using LangGraph to structure multi-stage tool-augmented reasoning. Without human intervention post-launch, selected open-source LLMs were tasked with scanning, vulnerability analysis, and exploitation against a standard testbed. Results show that models such as Qwen-14B and Qwen-32B can successfully execute multiple real-world exploits, demonstrating that local, API-free LLM agents can move beyond advisory roles into operational offensive security.
| Original language | English |
|---|---|
| Title of host publication | Recent advances in large language models (LLMs) have opened new opportunities for automating complex tasks in Cybersecurity, including offensive operations. However, most existing approaches to LLM-assisted penetration testing rely on human input or scripted interactions. This work explores a fully autonomous, local-agent framework for end-to-end penetration testing, using LangGraph to structure multi-stage tool-augmented reasoning. Without human intervention post-launch, selected open-source LLMs were tasked with scanning, vulnerability analysis, and exploitation against a standard testbed. Results show that models such as Qwen-14B and Qwen-32B can successfully execute multiple real-world exploits, demonstrating that local, API-free LLM agents can move beyond advisory roles into operational offensive security. |
| Publisher | Open Review |
| Number of pages | 9 |
| Publication status | Published - 3 Jun 2026 |
| Event | The Eighteenth Workshop on Adaptive and Learning Agents: ALA 2026 - Paphos, Cyprus Duration: 1 Jan 2026 → … |
Conference
| Conference | The Eighteenth Workshop on Adaptive and Learning Agents |
|---|---|
| Country/Territory | Cyprus |
| City | Paphos |
| Period | 1/01/26 → … |
Keywords
- LLMs
- Agents
- Penetration Testing
- Cybersecurity
- Agentic AI
- Lang- Graph
- Large Language Models
- GAAI
Fingerprint
Dive into the research topics of 'Let it hack: autonomous multi-agent penetration testing with LLMs and tool-augmented reasoning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver