Track 2 · Trust
Trust is not something the AI has. It's something your process produces.
"Can I trust AI?" is the wrong question. The right one is smaller and always answerable: how do I check this answer, given what it costs me to be wrong?
On this page: how hard to check · six checking moves · what isn't checking · what never to paste · failure modes · the checklist
Start here
The one fact that makes checking non-optional
A person who is unsure usually sounds unsure. They hedge, they slow down, they say "I think". An AI writes a fabricated legal case, a made-up statistic, and a correct historical date in exactly the same tone — fluent, calm, well-organised. The signal you have used your whole life to gauge reliability simply isn't there.
That is not a flaw someone forgot to fix. The system produces text that looks like the right answer. When it has the material, looking right and being right coincide. When it doesn't, only the first one survives — and it survives at full confidence.
Fluent, plausible, and impossible for you to check. A wrong answer you can test is a nuisance — you test it, it fails, you move on. A wrong answer about something outside your knowledge, delivered beautifully, goes straight into your work and your beliefs and stays there. Whenever you notice you're impressed and unable to verify, that's the moment to slow down.
Rule one
Match the checking to the cost of being wrong
Checking everything is exhausting; checking nothing is how people get hurt. Decide by what happens if this is wrong — before you read the answer, not after.
| If it's wrong… | Examples | How much to check |
|---|---|---|
| Nothing happens | Brainstorming, rewording your own sentence, ideas you'll judge yourself, explaining a concept you'll test in practice anyway. | Skim it. Your own judgement is the check. Move fast. |
| You look wrong | A report to your boss, a client email, a statistic you're about to post or repost publicly, code you're about to run, numbers in a deck. | Check every specific: each number, name, date, quote and claim. Run the code. Assume nothing is right because the rest was. |
| Someone gets hurt | Health, medication, money, legal rights, contracts, safety, tax, anything public, anything you can't undo. | An independent source or a qualified human. Always. The AI is for preparing your questions, never for the final answer. |
"If this is wrong, what happens, and who does it happen to?" Ten seconds. It moves the decision off how confident you feel and onto what is actually at stake.
Rule two
Six verification moves
Concrete and mostly fast. Pick by the triage above; make Move 4 automatic.
01Ask for sources — then actually open them
A citation from an AI is a claim that a source exists. It is not evidence. Systems that don't genuinely search the web can produce references with a real-sounding author, a real-sounding journal, and a title that was never written. Lawyers in several countries have been fined or formally reprimanded for filing invented case citations — the court checked, and the cases did not exist.
The move: ask for the source, then click it. If there's no link, search the exact title. If it doesn't exist, discard everything that rested on it — and be sceptical of the rest of the answer too.
The common version of this burn: reposting a tidy statistic, and three days later someone in the comments asks for the source. There isn't one.
Sources with no links, or links you can't open, are the ones to check first. A tool that genuinely searched will normally hand you something clickable.
02Ask again cold — and in a second model
Open a brand-new conversation and ask the same question without any of your earlier framing. Then ask a different AI product entirely.
- Same answer twice — mild encouragement, not proof. Two systems trained on overlapping text can be confidently wrong in the same way.
- Different answers — strong signal. You've just located exactly which claim needs a real source, for about thirty seconds of effort.
The "cold" part matters. Inside the same conversation it is anchored on everything already said, including your assumptions. A fresh chat drops that.
03Make it argue against itself
Paste the answer back and ask it to attack it. This works far better than asking "are you sure?", because you're giving it a different job rather than a chance to reassure you.
The list of separate claims is the real prize. Prose hides how many assertions you were handed; a list makes them countable and checkable.
04Check one thing you can check
The highest value per second on this page. In any answer there is usually something verifiable in under a minute: a total that should equal the sum of the parts, a date, a name, one link, one line of code.
Check exactly one. If it's wrong, stop trusting the whole answer — not just that line. Sloppiness in the checkable part is your best available evidence about the uncheckable part.
Never accept a set of numbers without adding them up once yourself. Silent arithmetic errors inside a well-formatted table are one of the most common ways AI mistakes reach a real audience.
05Prefer questions whose answers can be tested
You often get to choose the shape of what you ask for. Choose the testable shape.
| Instead of | Ask for |
|---|---|
| "Is this contract clause fair?" | "What does this clause let the other side do that I might not expect? List each one." |
| "What's the best way to do X?" | "Give me three approaches with the trade-off of each, and what would make me pick one over another." |
| "Write me a script that does X." | "Write it, plus a tiny test I can run on sample data to prove it worked." |
| "Summarise this." | "Summarise it, then list what you left out and why." |
Each right-hand version converts an opinion you'd have to trust into a set of statements you can inspect.
06Calibrate on ground you own
Ask it about the thing you know best in the world — your job, your hobby, your home town, the subject you studied. Read carefully. Count the errors, including the subtle ones a non-expert would sail past.
That number becomes your working assumption — your rough expectation of how often it's wrong — for every topic you can't check. Whatever you find, your own figure is worth more than any general claim about AI accuracy, including the ones on this page, because you measured it.
Repeat it when you switch tools, and once or twice a year. It moves.
Rule three
Four things that feel like checking but aren't
Almost useless. Push and it will often apologise and change its answer whether or not it was wrong — sometimes replacing a correct answer with a worse one. You learn about its agreeableness, not about the fact.
It will give you a number, produced the same way the answer was. It is a weak hint at best: a low number is worth heeding, a high number proves nothing. Never let it stand in for an actual check.
"I verified it" is a sentence, not an event. Re-reading its own answer can catch real errors — that is why Move 3 works — but a bare claim to have checked, with nothing shown, is not evidence. Look for the working or the tool output, not the reassurance.
The weakest signal of all. It leans towards agreeing with the person typing. If it confirms what you already believed, you have learned almost nothing — that's the moment to go looking for the counter-case.
The other half of trust
Trusting it with your information is a separate decision
Everything above is about trusting what comes out. This is about what goes in — a different question, with a different answer, and one people get wrong more casually.
Never paste
- Passwords, keys, one-time codes, card numbers
- Other people's personal or medical details
- Customer data, or anything under a confidentiality agreement
- Unreleased financials or anything price-sensitive
- Anything you'd be uncomfortable seeing quoted back at you
Check before you rely on it
- Whether your conversations may be used to train the model — it varies by product and by plan, and the setting can move
- Whether your employer has an approved tool, and whether the free consumer version is allowed
- Whether the account is personal or company-issued — that usually decides who can read it
- The provider's data settings page, in your own account, today
Before pasting, ask: would I be comfortable if this text appeared in a screenshot in a news story next year? If not, don't paste it. Removing names and numbers helps, but it does not make a document safe — a paragraph can still identify someone from context, and a redacted contract is still under the confidentiality agreement. When it genuinely matters, use the tool your employer has approved, or don't use one. Asking takes a minute. A leaked document can cost a career.
Pattern recognition
Eight failure modes, and how each one looks
Learn these by name and you'll start spotting them mid-read.
| Failure mode | Looks like | Catch it |
|---|---|---|
| Invented sources | A confident citation with no link, or a dead one. | Search the exact title. Thirty seconds. |
| Code that runs and is still wrong | Clean, no errors — and using the wrong column or date range. | Run it on an example whose answer you know. "No error" is not "correct". |
| Stale facts in the present tense | Present-tense claims about prices, laws or versions, frozen at the cutoff. | Ask how current it is, then check a live source. |
| Sycophancy | Eager agreement, and a plan with no downsides. | Ask for the case against. A real one has real costs in it. |
| Anchoring on your framing | You asked why X beats Y, and got an essay on why X beats Y. | Ask neutrally in a fresh chat, naming neither as the favourite. |
| Confident middle, invented edges | The explanation is solid; the exact figure or date is invented. | Trust the shape, check the specifics. Precision is where invention hides. |
| Silent unit and currency errors | Millions mixed with billions, percent with percentage points — in one tidy table. | Check the units are stated, consistent, and that totals add up. |
| Over-generalising from your one example | You gave one sample row; it works only for that row. | Test the awkward cases: empty, duplicate, missing field. |
A table, not eight fold-outs: this is a page people scan under time pressure, and nobody opens eight fold-outs.
Keep this
The one-page check
Run this on anything that leaves your desk. It takes about two minutes once it's a habit. Ticks are saved on this device.
The moment it goes out with your name on it, it is your work and your mistake. "The AI said so" has not saved anyone yet, and it is not something to plan around. That isn't a reason to avoid the tool — it's the reason to check it.