Test a Resume Prompt Before You Trust It With Your History
You cannot tell whether a resume prompt is any good by reading the prompt. Thoroughness on the page and thoroughness in the result are unrelated — the longest, most professional-sounding prompts often invite the most invention, because most of their length is a request for impact. The only way to judge one is empirically: run it on a piece of your history whose truth you already know cold, then examine the output for a short list of specific defects. This is a test procedure, not a prompt library.
Build one test input and reuse it forever
Write a single block of notes about one real job you remember clearly, and keep it in a file. This is your test rig, and it should be deliberately awkward, because you want to see how a prompt behaves under the conditions that actually break things. Include, on purpose:
- A thin line. Something genuinely vague that you cannot make more specific, because you did a small unmemorable thing. This is bait.
- A shared result. Something your team achieved where your part was real but partial.
- An unquantified outcome. A thing that improved, where you never measured by how much.
- A piece of internal vocabulary. A system name or process name that means nothing outside that employer.
- One hard fact you know exactly. A date, a headcount, or a frequency you are certain of.
Because you wrote it, you know the correct answer for every line. That is the whole point: a candidate prompt is now something you can mark, rather than something you can only admire.
Score the output on five defects
Paste your test notes under the candidate prompt, run it once, and read the result against the notes rather than as prose. Five things are worth counting.
1. Arrivals. Anything present in the output and absent from your notes. Numbers are the obvious case, but scope words are the frequent one — cross-functional, enterprise, end-to-end, company-wide. Every arrival is a defect regardless of how plausible it is, and plausibility is precisely what makes them dangerous later.
2. Verb escalation. Check what happened to your bait lines specifically. Helped becoming led, supported becoming owned, contributed to becoming delivered. A verb is a statement about your position relative to other people, which is why the verb you choose counts as a claim and not as style.
3. Ownership collapse. Look at what the prompt did with the shared result. Good behaviour keeps the distinction between what the team achieved and what you did inside it. Bad behaviour deletes the team, and the deletion is invisible unless you are comparing against your own notes.
4. Specificity loss. Compare the hard fact you were certain of. Did it survive, or did it become regularly or a large team? A prompt that rounds your one verifiable detail into a vague word has traded the most useful sentence in the batch for a smoother one.
5. Silence on the gap. This is the discriminating test, and it is the one most prompts fail. The thin line you planted has no good bullet in it. A well-constructed prompt notices and says so — asks you what the project was, or flags the line as unsupportable. A poorly constructed one produces a confident bullet, and the confidence is manufactured out of nothing.
The pass condition, stated plainly
A prompt passes if the output contains zero arrivals, zero escalated verbs, your hard fact intact, and an explicit note about the thin line. Anything else is a fail, and a fail is not fixable by trying again — it is a property of the instruction.
Note what is not on the list. Whether the bullets sound impressive. Whether they use industry language. Whether they are the right length. Those are cosmetic and you can adjust them in ten seconds. The defects above are the ones that produce a document you cannot defend in a conversation, and no amount of polish reaches them.
Why long, professional-sounding prompts often score worst
Read a few circulating resume prompts with the five defects in mind and a pattern shows up. The ambitious ones ask the model to highlight achievements, quantify results where possible, emphasise impact, and make the candidate stand out. Every one of those phrases is an instruction to add content, and if the notes do not contain achievements or numbers, the instruction can only be satisfied by inventing them. The prompt is not being careless; it is doing exactly what it says, and what it says is wrong.
The prompts that score well tend to be shorter and largely negative: use only what I gave you, add nothing, tell me what is missing. They read as unambitious and they behave as reliable.
Two mechanical notes on running the test
Run each candidate in a clean window. Testing three prompts in one conversation contaminates the results, because the second and third inherit the register of the first. The session-level reasons for this are in driving an AI session turn by turn, and they apply doubly when the thing you are measuring is the output itself.
Run the winner twice. Language models are not deterministic, and a prompt that passes once may pass because of luck. Two clean passes on the same input is a reasonable bar for something you are about to use on your actual history.
What a passing prompt does not buy you
A prompt that survives this test is safe to use, not safe to trust unread. It will still occasionally add something, and your notes will still be the only reference for whether a line is true — which is why every line that comes out of a passing prompt still gets checked against your own record before it reaches the document.
And a passing prompt cannot make a judgement call. It does not know which employer you are writing for, what to leave out, or whether a gap needs explaining. The test above measures one thing: whether the instruction, applied to real material, returns your history rather than a better-sounding version of somebody’s. That is a narrow property, and it is the only one worth selecting on.