How I Get AI to Tell Me When It's Guessing
The worst AI output I ever shipped was not wrong. It was 90% right, and the 10% was delivered in exactly the same confident register as the rest, so I had no way to see the seam. It was a paragraph about a platform's pricing tiers in a client deck. Three tiers were correct. One had been quietly invented, because the model had four slots to fill and three facts to fill them with.
Nobody caught it in the meeting. I caught it a week later and felt sick about it.
What I learned from that is that the standard advice — "always verify AI output" — is useless as stated. Verify *what*? If everything reads equally certain, verifying means re-researching the entire thing from scratch, at which point the model saved me nothing. The advice is only actionable if the output tells you where to look.
So that became the thing I actually prompt for. Not accuracy — I can't make a model accurate. Legibility about its own accuracy. Those are very different asks, and only one of them is achievable.
Why confident invention is the default
A model's job, as it experiences it, is to produce the most plausible continuation of your text. "I don't know" is almost never the most plausible continuation, because the training data is full of people confidently answering things and comparatively empty of people trailing off mid-sentence. Uncertainty is stylistically underrepresented.
Worse, hedging and inventing feel identical from the inside of the generation process. There's no internal flag that trips when it crosses from recall into construction. It's the same operation either way: what word probably comes next. Which means the smooth, authoritative tone you're reading is not evidence of anything — it's the house style, applied uniformly to things it knows cold and things it's assembling from fragments.
Once I stopped reading fluency as a signal, the practical question got much simpler: can I make it label the difference? Mostly, yes. Not perfectly — a model's self-assessment is itself a generation, not a readout of some internal certainty meter. But it's far better than nothing, and it's directionally honest in a way that surprised me. When I ask it to mark what it's least sure about, the things it marks are, in my experience, disproportionately the things that turn out to be wrong.
The prompt I use on anything factual
This is the one that goes in front of any answer I might repeat to another human.
Answer the question below. Then, underneath the answer, add a section called CONFIDENCE. In it, list every specific factual claim you made — numbers, names, dates, prices, quotes, feature availability — and mark each one: - SOLID: widely documented, you'd expect to find this stated identically in multiple sources. - LIKELY: you believe it but it may be outdated or you're reconstructing it from related facts. - SHAKY: you're inferring or filling a gap. Say what you inferred it from. If any claim is SHAKY, say explicitly what I'd need to check and where. Don't smooth this over — I'd rather have a short answer with three SOLID facts than a long one where I can't tell which parts to trust. Question: [question]
The last sentence does more work than it looks like. Without it, the model optimises for a complete-looking answer and grades generously — everything comes back SOLID because a confident answer feels like a better answer. Telling it that I prefer the short honest version changes what it's optimising for, and the gradings get noticeably harsher.
What I get back is a map of where to spend my checking time. Usually two or three items, which takes four minutes instead of forty. That ratio is the whole point.
The one that catches the invented fourth tier
The failure that burned me wasn't a wrong fact. It was a *structural* pressure toward invention: I'd asked for something in a shape, and the shape had more slots than reality had contents. Lists, comparison tables, "the five key X" — every one of these quietly asks the model to produce a fixed count, and it will hit the count.
So I prompt against the shape.
Give me [the list / table / comparison]. But: the structure is not a quota. If you only have four real items, give me four and say so. If a cell in the table has no reliable value, write UNKNOWN rather than a plausible one. Before the output, tell me in one line how many items you're confident actually exist, versus how many I asked for.
That one-line preamble has saved me more embarrassment than anything else in this post. "You asked for six; I can support four." Fine. Four is what I needed — I just didn't know that when I asked for six.
This applies well past facts, incidentally. Ask for ten content ideas and you'll get ten, of which four are good and six are padding written to reach ten. Same mechanism, lower stakes.
Auditing something already written
Half the time the problem isn't a fresh answer — it's a draft I already have, mine or the model's, that I'm about to send. For that I don't ask for a rewrite. I ask for a risk pass.
Read the text below as a hostile fact-checker who wants to find one thing in it that isn't true. Extract every checkable claim into a list. For each: what would have to be true for this to hold, and how would someone verify it in under two minutes? Flag anything that's stated as fact but is actually an interpretation, an extrapolation, or a number that sounds suspiciously round. Don't fix anything. Just tell me what I'd be sorry about. [text]
"What I'd be sorry about" is the phrasing that made this useful. Asking it to "check for accuracy" gets a polite pass. Asking what I'd regret makes it read like someone who's going to be quoted, which is the exact frame I need before a client sees something.
What this doesn't fix
Being straight about the limits: a model can be confidently wrong about its own confidence. I've had SOLID-marked claims turn out stale, because the fact was true and then the world changed. The confidence pass is a prioritisation tool, not a guarantee. It tells you where to look first; it doesn't replace looking.
And there's a category it simply can't help with: things the model has never encountered but that sit in well-trodden territory, so the invention comes out fluent and structurally perfect. Small companies' pricing, recent API changes, anything in the last few months. For those, no prompt saves you. You check, or you don't claim it.
What the confidence pass buys is the ability to *know that you're in that category*. The failure mode I actually needed to kill wasn't being wrong — it was being wrong without knowing there was a question. That one, this fixes.
The habit, not the prompt
None of these are clever. Their entire value is that they run every time, on the boring middle of the work, not just on the parts where I already suspect something. The invented pricing tier got through precisely because it was in a paragraph I wasn't worried about.
Which means it can't depend on remembering. I want the answer, and appending a hundred words about grading your own certainty is the last thing I feel like typing when I want the answer. So it lives one click away, the same as everything else I know works and would never type in the moment. That's what Super Prompts actually is for me — not a library I browse, a set of reflexes I don't have to have.
Try the first one on the next thing you'd repeat to someone else. Read the SHAKY items. In my experience there's usually one, it's usually a number, and you'd usually have sent it.
If you want prompts that make AI show its work instead of just sounding sure, keep them where reaching for them costs nothing — Super Prompts is free to start.