Here are the AI use cases people in my organisation have proposed: [list them].<br/><br/>Score each on four things, and show your reasoning:<br/>1. Is the current process well enough defined that we could tell whether AI did it correctly?<br/>2. Is there a named person whose job gets measurably easier?<br/>3. What happens when it is wrong - is that recoverable or not?<br/>4. Could we ship something useful in six weeks, or is this a year?<br/><br/>Then tell me which single one to start with, and which to refuse outright even though someone senior is attached to it.
Help me write the success criteria for this AI deployment before I build it: [describe the use case].<br/><br/>I want numbers I can measure in production, not a vision statement. For each, say what the current baseline is, what good looks like, and how we would collect it.<br/><br/>Include at least one criterion that would make us switch it off. If I cannot name the condition under which we abandon this, I am not running a project, I am running a demo.<br/><br/>CURRENT PROCESS:<br/>[how it works today, with any numbers you have]
I am taking this AI deployment to security review: [describe the system, what data it touches, what it can do].<br/><br/>Play the security reviewer. Ask me the hardest questions you would ask, in the order you would ask them - data handling, retention, access, what the model can act on, what happens on failure, what gets logged, how we would detect misuse, vendor terms.<br/><br/>After each question, tell me what a weak answer looks like and what a sufficient one looks like.<br/><br/>Do not go easy. I would rather be embarrassed here.
Help me design a pilot for this AI deployment that will produce a decision rather than an anecdote: [describe the use case and success criteria].<br/><br/>Tell me: who is in it and who is deliberately not, how long it runs, what we measure, what the comparison is, and how many cases we need before the result means anything.<br/><br/>Then name the ways this pilot could produce a flattering result that does not generalise - self-selected enthusiastic users, cherry-picked inputs, me being available to fix things in a way nobody will be at scale.<br/><br/>Tell me how to design those out.
For this deployment: [describe it], map every way the output can be wrong, and for each one tell me:<br/>- how someone would notice<br/>- how long it would take them to notice<br/>- what the damage is in that window<br/>- who is accountable<br/><br/>Separate the failures that are loud and obvious from the ones that look like correct output. The quiet ones are the dangerous ones.<br/><br/>Then tell me what to put in place for the quiet ones specifically - what to log, what to sample, what to review and how often.
I am handing this AI deployment to a team that did not build it: [describe the system and who is receiving it].<br/><br/>Write me the handover document. It needs to cover: what it does and deliberately does not do, how to tell it is working, the three most likely failure modes and what to do about each, what to check when the vendor ships a model update, who to contact, and what decisions were made and why.<br/><br/>That last section matters most. Write it so the next person does not redo an experiment I already ran.<br/><br/>Flag anything I have not told you that the document needs.