goJumboGPT

AI AI tools compared: free versus paid, search and phones

How to choose an AI assistant without reading a benchmark

A durable way to compare AI assistants: the six questions that matter more than benchmark scores, and how to run a fair test with your own work in an afternoon.

8 min read How we write

The short answer

  • Choose an AI assistant by running ten of your own real tasks on two or three candidates and scoring the answers blind, not by comparing benchmark scores.
  • Six criteria settle it: task fit, how much material it holds at once, what it can reach, the data terms, the shape of the price, and how easily you could leave.
  • Benchmark gaps between the leading assistants are usually smaller than the difference a better prompt makes, which is why they rarely decide a real choice.
  • A headline context length is a ceiling rather than a promise, so test with your own longest document and ask a question about its middle.
  • When two come out close, decide on data terms and price shape, because those are contractual and will still be true in a year.
  • Keeping a second free account on a different assistant costs nothing and gives you a way to sanity check an answer before you act on it.

Choose on fit, not on scores. Collect ten tasks you actually do, run them on two or three assistants, hide which answer came from which, and pick the one that wins on your own work. A leaderboard ranks models on averaged test questions. It cannot tell you which one keeps your table formatting, reads your scanned invoices, or writes in a register your clients will accept. Six things decide it in practice: task fit, how much material the tool holds at once, what it can reach, what the provider does with your data, the shape of the price, and how hard it would be to leave.

Why benchmark scores do not settle it

A benchmark is a fixed set of questions with known answers. The number is real, but three things make it a poor buying guide. The questions are not your questions: competition math and graduate level science separate frontier models from each other and say nothing about turning a messy transcript into an action list. Scores compress: when several assistants sit a point or two apart, that gap is smaller than the difference a better prompt makes. And benchmark questions leak into training data over time, so a rising score can partly reflect familiarity with the test rather than skill at the task.

There is a fourth problem that has nothing to do with testing. You are not buying a model, you are buying a product built around one. File handling, export, search and the behavior when you paste sixty pages in are all part of what you touch daily. Two assistants running similarly capable models can feel completely different because one silently trims your document to fit and the other tells you it did. That difference matters more than a benchmark point, so read what happens when a document exceeds the context window before you test anything long.

The six questions that decide it

Ask these in order. The first three are about capability, the last three about consequences. The last three are the ones people skip.

Task fit

Not whether it is smart, but whether it is good at the shape of work you bring. Assistants differ by category: long form drafting, code, structured extraction from documents, translation, arithmetic heavy analysis, image work, conversation in your language. A model that writes graceful prose can be mediocre at returning clean, machine readable output, and the reverse happens too. This is the one criterion you cannot outsource to somebody else's review, because nobody else has your work.

Working memory

How much can it hold in view at once, and what happens when you exceed that? A headline context number is a ceiling, not a promise. Recall usually weakens toward the middle of very long inputs, and some products cap uploads well below the model's own limit. Test with your longest document, then ask a question whose answer sits halfway through it.

Reach

What can it touch beyond the chat box? File uploads and the formats it truly parses (a scanned PDF is a picture, not text), live web access, code execution, connections to mail and storage, and the ability to take actions for you. Most of the real time saving lives here rather than in raw model quality, and so does most of the risk.

Data terms

Whether your inputs train future models by default, how long conversations are kept, who can read them for abuse review, and whether consumer terms differ from business ones. If you handle anyone else's personal data, this outranks answer quality. It is worth knowing where your prompts actually go after you press send, because for some organizations the processing region matters as much as retention.

Price shape

Not the headline number, the shape. A flat subscription with generous caps, a credit pool that drains faster on images and long documents, and per token billing behave very differently at the same nominal cost. The three pricing models and how they add up is worth reading before you commit to anything annual.

Exit

How much work would leaving be in a year? Check whether chats export in a usable format, whether saved prompts and custom instructions are portable, and whether anything you build (custom assistants, shared workspaces, automations) lives only inside that product. Prompts travel anywhere. Configuration does not.

What to look for, criterion by criterion

Keep this open while you run two tabs side by side. The right hand column is why each check earns its five minutes.

CriterionWhat good looks likeHow to test it in five minutesThe failure it prevents
Task fitWins on your own work, not a demo promptRun your three hardest real tasksMonths of small daily friction
Working memoryAccurate about the middle of a long filePaste your longest document, ask about its middleA confident summary that quietly skips a section
File handlingReads scans and spreadsheets cleanlyUpload the messiest file you ownRetyping the data by hand anyway
ReachReal web access, with sources you can openAsk about something from this monthFluent answers built on stale knowledge
Data termsTraining off by default, retention stated, business tier availableFind and read the data controls pageClient material sitting in a training set
Price shapePredictable at your volume, with visible limitsRead the limits page, not the pricing pageRunning out halfway through a deadline
ReliabilityConsistent during your working hoursUse it for a week at your normal timesQuietly routed to a weaker model at peak
ExitOne click export of chats and saved promptsRun an export on day oneA year of work locked in one product

Build an eval set of ten real tasks

A personal eval set is ten tasks, written once and reused. It takes about an hour to assemble and it outlives every model release, because your work changes far more slowly than the models do.

  1. Go through last month and pull ten things you actually asked for, or wished you could: the summary of a particular report, the rewrite of a particular email, the extraction you did by hand.
  2. Include two tasks where you already know the correct answer. These are your lie detectors, and the only way to catch an assistant that is fluent and wrong.
  3. Include one task at the edge of what you expect to work: a hundred page document, an ambiguous question, a file format that usually breaks things.
  4. Write down what a good answer looks like before you run anything. Three bullet points is enough. Deciding afterward is how you talk yourself into the assistant you already liked.
  5. Freeze the wording. The same prompt goes to every candidate, or you are testing your prompting rather than the tools. Weak prompts make the whole comparison noise, so fix the parts of a prompt that carry the weight first.
  6. Save the set where you will find it again. This is also the start of a prompt library you will actually reuse.

Score it blind

Brand halo is real and strong. If you know which answer came from the assistant you already pay for, you will rate it higher. Order matters too: whichever answer you read first becomes the standard you judge the rest against, which is why the second one so often reads like a rearrangement of the first.

The fix is mechanical. Paste each answer into a document labeled A, B and C, shuffle the labels, keep the key elsewhere, and leave it overnight if you can. Then score on three things only: did it do the task, is it right where you can check, and how much editing would it need. A three point scale beats a ten point one, because ten point scales collapse into giving everything a seven.

Count wins per task rather than averaging. Averages hide the case that matters: an assistant slightly better on eight tasks and unusable on two is the wrong choice when those two are weekly work.

The tie breakers when two are close

A close result is the normal outcome, not a failed experiment. Decide it on the things that do not change week to week.

  • Data terms and processing region, which are contractual and slow to move.
  • Price shape at your heaviest day rather than your average one, since that is when you will hit a wall.
  • Whether most of your questions are lookups rather than conversations. If so, your real decision is between an AI answer and a search engine, not between two chatbots.
  • Whether some work should never leave the machine at all. For that slice, the option is running a model on your own hardware alongside a hosted assistant for everything else.
  • Whether the mobile app, the keyboard shortcut and the export fit how you already work. Small friction compounds daily.

The afternoon that settles it

Open free accounts on two or three assistants. Spend forty minutes assembling your ten tasks and writing down what a good answer looks like. Run them, paste the results into a blind document, and score them the next morning. Then check the two things no test can show you: the data controls page and the usage limits page. If you are still unsure whether to pay at all, what a paid tier actually adds over a free one answers that separately, and for occasional users the answer is usually no.

Rerun the set when something significant changes, which in practice means once or twice a year rather than every time a version number moves. Keep a second free account on a different assistant permanently. It costs nothing and gives you somewhere to check an answer that looks a little too confident before you act on it.

Common questions

Which AI assistant is the best right now?

There is no stable answer, and that is the useful part. The lead changes with every release cycle, and the gap between the top few is small compared with the gap between a vague prompt and a good one. Ask instead which is best at your ten tasks, on terms your data can live with. That answer changes far more slowly.

How many assistants should I actually test?

Two or three. Beyond that, scoring fatigue sets in and the later answers get judged more harshly than the early ones, so the test stops measuring the tools. Pick candidates that differ in something that matters to you, such as file handling or data terms, rather than three variations on the same product.

Are AI leaderboards completely useless then?

No. They are good at spotting a generation gap, for example when one model is clearly behind the pack rather than a point or two off. They are poor at ranking close rivals, because the questions are narrow, the margins are thin and test material tends to leak into training data over time. Use them to build a shortlist, not to pick a winner.

How often should I redo the comparison?

Once or twice a year, or when your own work changes shape. Because you froze the prompts, rerunning the set takes under an hour. Resist rerunning it every time a new version is announced, since most releases move your ten tasks very little and the switching cost is real.

Can I just use two assistants at the same time?

Many people do, and it works well if one is paid and one is a free account used for cross checking. Paying for two rarely makes sense, because the second subscription buys capacity you will not use. The exception is when your work splits cleanly, for example one tool for code and another for long documents.