Track 4 · Tools
Don't learn the leaderboard. Learn the categories.
New model names arrive weekly, each announced as the best; memorising them is a treadmill. The categories underneath barely move — know your task's category and picking a product takes a minute.
Anything specific written here would be wrong within months, and you'd have no way to tell which parts had rotted. So this page teaches the shape of the market and how to check the current state yourself, which stays true. Written 5 August 2026.
The six shapes
Six categories. Almost every AI product is one of these.
Vendors blur the lines and bundle several into one app. That's fine — you still need to know which one you're asking for, because it decides what "good" means.
01General chat — the everyday one
The one most people mean by "AI". Fast, conversational, good at writing, explaining, summarising and reformatting. This handles the large majority of ordinary use.
Pick it for: writing, explaining, drafting, tidying, translating, everyday questions.
02Reasoning — the slow, careful one
A mode or model that spends noticeably longer before answering, working through the problem in steps. Better on multi-step logic, tricky analysis, maths and debugging. Slower, and usually more expensive or capped at fewer uses per day.
Pick it for: problems where the answer depends on getting several steps right in order. Don't use it to reword an email — you'll wait longer for no benefit.
03Search-connected — the one that reads the web now
Looks things up live and cites what it found, instead of relying only on training data. Essential for anything current: prices, news, laws, product comparisons, who-currently-holds-what.
Its value is the links. Open them. A search-connected answer you didn't click through is no more verified than any other — see Trust.
04Agents and coding tools — the ones that act
Take multiple steps without asking each time: search, run code, edit files, check their own output, retry. The biggest jump in capability, and the biggest jump in what can go wrong unattended.
Pick it for: real multi-step work — building something, processing many files, anything that would take you an afternoon of repetitive steps.
05Media — images, audio, video, voice
Generating or editing pictures, transcribing recordings, reading text aloud, holding a spoken conversation, producing video. Moving faster than any other category, so nothing specific is worth memorising.
Who owns what you generate varies by product and by plan — check before using it commercially. And generated people, voices and events can be indistinguishable from real ones, which is now your problem as a consumer of media too, not just as a maker.
06Local / open-weights — the one that runs on your machine
Models you download and run on your own computer — "open-weights" simply means the model file is yours to download. Typically less capable than the best hosted ones, but nothing is sent to a company's server, there is no subscription (you pay in hardware and electricity instead), and it works offline.
Pick it for: genuinely sensitive material, or curiosity. Needs a reasonably powerful computer and some patience. Not the place to start.
The decision
Which one, in four questions
Is the material genuinely confidential?
Yes → your employer's approved tool, or a local model, or redact before pasting. Ask this before the other three: the rest are reversible, and this one is not.
Does the answer depend on something recent?
Yes → search-connected, and open the links. No → carry on.
Does it need several steps to be right in sequence?
Yes → the reasoning mode, and expect to wait. No → general chat is faster and usually just as good.
Does something need to be done, not just written?
Yes → an agent or coding tool, working on copies, with you reading what it did. No → stay in chat.
Within a category, the products from the major labs are similar enough that your context, your iteration and your checking usually matter more than the choice. Don't take that on trust either — it is what your own five-prompt test below is for. Pick one, learn it properly, and re-test each quarter. Switching every month costs more than it gains.
The thing nobody explains
"Free until you hit the limit" — what's actually going on
Free plans feel like a trick when the limit appears. Mostly they are not. Every answer costs the provider real money in electricity and hardware — unlike ordinary software, where one more user costs almost nothing. A free plan is a sample, sized to let you find out whether it's worth paying for.
What that means in practice
- Free usually means the smaller or faster model, or a limited number of the good answers per day. The gap between free and paid is often larger than the gap between two vendors.
- Limits reset on their own schedule, which varies by product and changes without notice. Your account page is the only current answer.
- The limits move without warning. Providers change them as capacity and costs change. If a number is written somewhere, treat it as a snapshot.
- Whether your conversations are used for training varies by product and by plan, and often differs between free and paid. It is a setting in your own account — go and read it.
- "Unlimited" is rarely literally unlimited. There is almost always an upper bound for unusually heavy use, whether or not it is published.
If you use it a few times a week, free is genuinely fine — stay there. If you're hitting limits mid-task, the interruption is costing you more than the subscription. One paid plan, chosen properly and kept for a while, beats juggling three free accounts to dodge limits — which mostly teaches you three interfaces and no depth.
The durable skill
Build your own five-prompt test
Public benchmarks and leaderboards measure things that may have nothing to do with your work. The only ranking that matters to you is on your own tasks — and it takes about twenty minutes to build a test you can reuse for years.
Pick five real tasks from your own week
Real ones, with the real material. Cover different shapes: one writing, one explaining, one analysing something with numbers, one about your specialist subject, one awkward or messy task that tools have failed at before.
Write down what a good answer looks like — before you run it
Do this first or you'll be persuaded by whichever answer is best-formatted. One line per task is enough.
Run all five through each candidate, with identical wording
Same text, fresh conversation each time, no follow-ups. Any difference in how you asked contaminates the comparison.
Score them yourself, on your own criteria
Right / nearly / wrong is enough. Count the errors on your specialist task especially — that's the one where you can actually see the truth.
Keep the five prompts. Re-run them when something new launches.
Now every announcement is a twenty-minute check instead of an anxious guess. This is the single most useful thing on this page.
Public leaderboards move constantly and often measure exam-style questions. A model that tops a chart can still be worse at your actual job than one three places below. Your five prompts are worth more than any chart, because they're graded by the only person whose opinion decides whether you got value.
Staying current without drowning
How to check today's state of play, in five minutes
Since nothing specific on this page can stay accurate, here's the routine that replaces it. Once a quarter is plenty for most people.
- Go to the source, not the commentary. Each major lab publishes its own model list, limits and pricing. The provider's own documentation is the only thing that is current by definition.
- Read your own account page. Your plan, your limits and your data-sharing setting are all there, and they're specific to you rather than to an article.
- Use a search-connected AI to summarise the change — then open the links it gives you. That's the whole point of that category.
- Re-run your five prompts. Not part of the five minutes — this one needs twenty of its own. It answers the only question you actually had.
- Ignore anything that ranks models without telling you what it tested. An unexplained ranking is an opinion in a table.
Judge an AI teacher by five things: do they show a task failing as well as succeeding; do they name the model and the date; do they teach checking, not just prompting; do they show the boring middle of the work rather than a cut to the result; and are they selling a course at the end. Two out of five is usually not worth an hour.