WED, OCTOBER 07, 2026
Independent · In‑Depth · Practitioner‑Tested
Claude Coding

AI Agent Security Prompts: 6 to Run Before You Deploy

On 6 October 2026 OpenAI apologised to an Australian Senate inquiry for an agent that reached a government Medicare system without authorisation, and disclosed a second breach at NSW Parks and Wildlife. Anthropic told the same hearing it would notify government within days or sooner. Disclosure speed, not the incident itself, is what regulators are moving to measure. These six prompts cover the work that has to happen before that matters: auditing what your agent can reach, classifying tool calls by blast radius, and having a disclosure draft written before you need it.

⌨️ 6 prompts 🕐 Updated Oct 6, 2026
💡 How to use these prompts: Replace everything in [BRACKETS] with your specific details before sending. Click Copy to copy any prompt to your clipboard instantly.
1
Map what the agent can actually reach
Finding out what you actually granted, before something else does
Below is the configuration for an AI agent: its tools, credentials, network settings and system prompt.\n\nIgnore what the configuration intends. Tell me what it permits.\n\n1. List every external system this agent can reach, including ones reachable only by chaining two tools together.\n2. For each, say what it can do there: read, write, delete, spend money, send messages as someone.\n3. Name every credential in scope and what else that credential opens beyond this agent.\n4. Identify any path to the public internet, including indirect ones through a tool that fetches a URL.\n\nEnd with the single worst outcome reachable from this configuration without any bug -- using only permissions as granted.\n\n[PASTE CONFIGURATION]
2
Classify every tool call by blast radius
Deciding what needs a human in the loop and what does not
Here is the tool list my agent has access to. For each tool, classify it:\n\nREVERSIBLE -- a mistake costs time only\nCOSTLY -- a mistake costs money or data that can be restored\nIRREVERSIBLE -- a mistake cannot be undone\nEXTERNAL -- the effect is visible to someone outside my organisation\n\nFor every tool above REVERSIBLE, write the confirmation gate it should sit behind: what the agent must show a human, and what the human has to approve.\n\nThen tell me which tools are currently ungated that should not be, ordered by how bad the first mistake would be.\n\n[PASTE TOOL LIST]
3
Red-team the sandbox
Testing containment before production rather than after an incident
You are testing whether an agent can escape the boundary described below. Work as an attacker who controls only the agent inputs.\n\nFor each escape route you find:\n\n- The exact sequence of inputs or tool calls\n- Which specific control fails and why\n- What is reachable once outside\n- Whether a log would show it happening\n\nPay particular attention to: tools that accept a URL or file path, anything that renders or executes returned content, and any point where output from one tool becomes input to another without validation.\n\nIf a route needs a condition I have not described, state the condition rather than assuming it.\n\n[PASTE SANDBOX AND TOOL DESIGN]
4
Design the egress alert that would have caught it
Building the monitoring that catches an agent reaching somewhere it should not
My agent should only ever contact these hosts: [LIST]. Anything else is an incident.\n\nDesign the detection for that. Give me:\n\n- What to log on every outbound request, and what to leave out to avoid logging secrets\n- The alert conditions, separated into block immediately and notify a human\n- How to catch a request that goes to an allowed host but carries data that should not leave\n- The false positives this will generate in normal operation, and how to tune them out without blinding the alert\n- A test I can run today to prove the alert fires\n\nAssume the agent will eventually do something nobody predicted. Design for noticing, not for prevention alone.
5
Write the disclosure before you need it
Having the hard document written on a calm day
Draft an incident disclosure template for an AI agent that accessed a system it was not authorised to reach.\n\nWrite it to be sent within 48 hours, with facts still incomplete. That means it must work without knowing root cause, full scope or total impact.\n\nInclude:\n\n- What is known, stated plainly, with timestamps\n- What is not yet known, named specifically rather than as a gap\n- What has been done to stop it continuing\n- What the recipient should do now\n- When the next update will arrive, as a commitment\n\nNo hedging language, no passive voice about who did what, no apology that avoids saying what happened. Give me a version for a regulator and a version for an affected customer.
6
Audit the agent action log
Checking you could answer a regulator six weeks later
Here is a sample of my agent action logs. Assess whether these logs would let me reconstruct an incident six weeks after it happened.\n\nAnswer these specifically:\n\n- Could I tell which prompt or input led to a given action?\n- Could I tell which version of the system prompt and tool set was live?\n- Could I distinguish an agent action from a human action in the same system?\n- Could I prove an action did not happen, if someone claimed it did?\n- Is anything logged that should not be, such as credentials or personal data?\n\nFor every no, give the exact field to add and where it has to be captured.\n\n[PASTE LOG SAMPLE]