Here is my system prompt, policy or config: [paste].<br/><br/>Find every clause that overrides, supersedes, excepts, or takes precedence over another clause. Quote each one exactly and tell me what it suspends.<br/><br/>Pay particular attention to words like unconditional, always, never, regardless, notwithstanding, overrides, takes priority, and any sentence beginning "unless".<br/><br/>Rank them by how much they undo. The one at the top is the one that actually governs the document.
Here is a clause from my policy: [paste it].<br/><br/>Tell me what it literally covers, as written, rather than what it was obviously meant to cover.<br/><br/>Then give me three situations that fall inside the literal wording and clearly outside the intent. Be specific and realistic, not contrived.<br/><br/>Finally, rewrite it so the wording matches the intent, and tell me what capability I lose by narrowing it.
I want to test what my deployed assistant will disclose about itself.<br/><br/>Give me a sequence of requests to try, starting with the most direct - asking plainly for its instructions, its configuration, its files - and escalating only to things a normal curious user would try.<br/><br/>For each, tell me what a correct refusal looks like and what a leak looks like.<br/><br/>No exploit chains. The point is to find out whether plain asking works, because that is what actually happened to Meta.
My assistant or agent has these permissions: [list them].<br/><br/>Work out what it can actually reach as a consequence - not what I granted, but what those grants transitively allow. Files, services, accounts, other people's data that arrives through mine.<br/><br/>Flag anything it can read about people who are not users and have not consented - contacts, message senders, shared documents, calendar invitees.<br/><br/>Tell me which permissions I could remove without losing the thing I actually use it for.
My system has a layer that is supposed to be non-overridable: [describe it - a guardrail, a policy engine, a permission boundary].<br/><br/>Ask me the questions needed to work out whether that is actually true: where it sits, what calls it, whether anything can run before it, whether any config can disable it, who can change it and with what approval.<br/><br/>Then tell me what evidence would demonstrate it holds, and whether I have that evidence or just an assurance.
Here is my complete policy, prompt or config: [paste].<br/><br/>You want to get it to do something it should refuse. Not through a clever exploit - through its own words.<br/><br/>Which clause would you invoke? What framing would you use? Which exception would you argue you fall under?<br/><br/>Show me the three strongest arguments available inside the document itself, and then tell me which edit closes each one.