About this project

AIxploit is an automated security testing framework designed to discover AI-native vulnerabilities in agentic systems. It is useful for offensive testing against a graybox agent and for model developers to measure and improve safety guardrails against real, end-to-end injection attacks. The framework addresses the danger of AI agents being given real tools (databases, file systems, shells, MCP servers) and acting on data they don't control. It automates the exploration of attack surfaces by having agents brainstorm injects, then fires them at a target inside a disposable sandbox, checking whether the agent was manipulated. It ships with three attack scenarios: - `ransom`: A support agent triaging tickets is talked into encrypting customer emails via SQL injection. - `postgres`: An agent on a read-only database is pushed to escape read-only mode and copy secrets. - `kyc`: An attack hidden inside a passport extraction task dumps other customers' records. How it works: - **Prepare**: Generate inject corpus using `prompt_explorer` (agents read summaries of existing prompts to invent new techniques). Set up disposable Docker containers, data, agent instructions, and custom routines for seeding and checking success. - **Run**: `aixploit.py` starts the sandbox, plants the inject, lets the agent run, and collects successful injects. - **Collect exploits**: Results are saved in `runs/` and confirmed hits in `exploits/`. Quick start: Install dependencies, configure API credentials, build prompts, and run `python aixploit.py --cfg cfg/ransom.yaml`. Prompt generation is automated via a Claude agent that reads only summaries of existing prompts to generate new ones along five dimensions: authority framing, precondition framing, compliance pressure, action specification, and technical mimicry. Users can add their own agent and test by providing templates, data, custom routines, instructions, prompts, and a config YAML. Warning: The included configurations contain hardcoded credentials for internal use only; do not use in production.