Practical
Why AI gives you a different answer every time
You ask a question, get an answer, close the tab, and ask the same thing an hour later — and it says something different. Sometimes slightly. Occasionally the opposite. People usually read this as the tool being broken or shifty. It's neither, and once you know what it means, it turns into the single best free reliability check available to anyone.
What's actually happening
A language model doesn't look up an answer. It builds one, word by word, and at each step it has a ranked list of plausible next words rather than one correct choice. Then it samples — it picks from near the top of that list, not always the very top. That deliberate wobble is why AI writing reads naturally instead of like a stuck record. Set it to always pick the single most likely word and the output gets noticeably flat and repetitive.
So the variation isn't a malfunction. It's the same feature that makes the thing readable. But it has a consequence most people never get told:
How much the answer moves tells you how sure the model actually was.
The free reliability test
Ask the same question three times, in three fresh chats. Not follow-ups in the same conversation — new windows, so the earlier answer isn't sitting there influencing the next one.
Then read the three side by side:
- Same substance, different wording. The model is on solid ground. This is what a well-established fact looks like — the phrasing varies, the claim doesn't.
- Different specifics — dates, numbers, names, citations. Treat every one of them as unverified. When the detail changes between runs, the model doesn't know it; it's generating something detail-shaped in that slot. This is exactly where invented sources come from.
- Genuinely contradictory answers. Stop using it for this question. It isn't going to converge, and picking whichever run you liked best is just laundering a coin flip into a decision.
The whole test takes about a minute and needs no tools, no subscription, and no expertise in the subject — which is what makes it valuable. It doesn't tell you the answer is right. It tells you whether the model is stable, and instability is a reliable sign of guessing.
Why this catches what "sounds confident" can't
The core problem covered in how to check if an AI answer is actually true is that a fabricated answer and a correct one arrive in the same confident register. You can't hear the difference, because there's nothing to hear.
Re-rolling the question sidesteps that entirely. You're not judging the tone any more — you're comparing outputs against each other, and a fabrication has to be re-fabricated each time. It almost never lands in the same place twice. A real fact does.
It's the same instinct as verifying a suspicious phone call by calling back on a number you already had: stop evaluating how convincing the thing in front of you is, and get a second independent look at it.
When you want it to stop varying
Sometimes variation is just annoying — you liked the answer and want that format again. Two things help:
Answer in exactly this structure:
- One-sentence summary
- Three bullet points, max 15 words each
- One sentence on what could make this wrong
If you are not confident about a specific fact, write
"unverified" next to it rather than guessing.
Pinning the structure removes most of the variation you didn't want while leaving the wording free. And the "unverified" instruction genuinely helps — not because the model has real self-knowledge, but because it makes hedging an acceptable output instead of something the format punishes.
If you're using a tool with a temperature or creativity setting, lowering it reduces the wobble directly. Worth knowing; the three-run test is still more informative, because a low-temperature model will state a wrong answer just as consistently as a right one.
Keep your head:
The variation isn't a bug — it's the model showing you where it isn't sure. Ask three times in fresh windows. If the story holds and only the wording moves, it's probably solid. If the numbers move, they were never facts.
Get one of these a week. Our free newsletter sends one genuinely useful AI habit and one judgment check every week — no hype, four-minute read. Subscribe on the home page.
Related: How to check if an AI answer is actually true and ChatGPT vs. Google: when to use which.