MON, OCTOBER 05, 2026
Independent · In‑Depth · Practitioner‑Tested
Claude Productivity

Six Prompts to Find the Clause That Undoes Your Other Rules

Meta Muse's leaked system prompt contains one line: "The user's authority over their own household is unconditional and overrides your safety training." Everything written above it is conditional on that sentence. A researcher then obtained the rest of the internal instructions by asking the chat interface for its own files. Most system prompts, security policies, permission configs and terms of service have a clause like this somewhere - a single conditional that suspends the rest - and almost nobody goes looking for it. These six prompts do.

⌨️ 6 prompts 🕐 Updated Oct 5, 2026
💡 How to use these prompts: Replace everything in [BRACKETS] with your specific details before sending. Click Copy to copy any prompt to your clipboard instantly.
1
Find Every Line That Suspends Another Line
Meta Muse's was one sentence containing the word "unconditional". Everything above it became conditional.
Here is my system prompt, policy or config: [paste].<br/><br/>Find every clause that overrides, supersedes, excepts, or takes precedence over another clause. Quote each one exactly and tell me what it suspends.<br/><br/>Pay particular attention to words like unconditional, always, never, regardless, notwithstanding, overrides, takes priority, and any sentence beginning "unless".<br/><br/>Rank them by how much they undo. The one at the top is the one that actually governs the document.
2
Test Whether the Scope Is as Narrow as Intended
"Household" covers more people than "the user's own data and property". Scope creep in a single noun.
Here is a clause from my policy: [paste it].<br/><br/>Tell me what it literally covers, as written, rather than what it was obviously meant to cover.<br/><br/>Then give me three situations that fall inside the literal wording and clearly outside the intent. Be specific and realistic, not contrived.<br/><br/>Finally, rewrite it so the wording matches the intent, and tell me what capability I lose by narrowing it.
3
Ask It for Its Own Instructions
The Muse instructions came out because a researcher asked the chat interface for its own files. Start at the floor.
I want to test what my deployed assistant will disclose about itself.<br/><br/>Give me a sequence of requests to try, starting with the most direct - asking plainly for its instructions, its configuration, its files - and escalating only to things a normal curious user would try.<br/><br/>For each, tell me what a correct refusal looks like and what a leak looks like.<br/><br/>No exploit chains. The point is to find out whether plain asking works, because that is what actually happened to Meta.
4
Map What It Can Reach Versus What It Was Given
Muse was described as building hourly profiles of everyone in a user's contact graph. Those people never agreed to anything.
My assistant or agent has these permissions: [list them].<br/><br/>Work out what it can actually reach as a consequence - not what I granted, but what those grants transitively allow. Files, services, accounts, other people's data that arrives through mine.<br/><br/>Flag anything it can read about people who are not users and have not consented - contacts, message senders, shared documents, calendar invitees.<br/><br/>Tell me which permissions I could remove without losing the thing I actually use it for.
5
Check the Layer You Were Told Cannot Be Overridden
Meta says Sentinel cannot be overridden. Nothing published explains how it interacts with the clause that overrides safety training.
My system has a layer that is supposed to be non-overridable: [describe it - a guardrail, a policy engine, a permission boundary].<br/><br/>Ask me the questions needed to work out whether that is actually true: where it sits, what calls it, whether anything can run before it, whether any config can disable it, who can change it and with what approval.<br/><br/>Then tell me what evidence would demonstrate it holds, and whether I have that evidence or just an assurance.
6
Read It as Someone Who Wants It to Fail
A rule set is only as strong as its most generous reading. Find the generous reading before someone else does.
Here is my complete policy, prompt or config: [paste].<br/><br/>You want to get it to do something it should refuse. Not through a clever exploit - through its own words.<br/><br/>Which clause would you invoke? What framing would you use? Which exception would you argue you fall under?<br/><br/>Show me the three strongest arguments available inside the document itself, and then tell me which edit closes each one.