Blog

How to Evaluate an AI: the One That Impresses Is Not the One That Gets It Right

Two models from the same lab were tested this week. The smaller one won the work benchmark — by presenting better, not by being more correct: the objective check came in at 20% versus 19%. And it scores negative on knowledge, getting more wrong than right. What that teaches about choosing AI for your company.

July 31, 2026 · Josué Gomes

How to Evaluate an AI: the One That Impresses Is Not the One That Gets It Right

This week, Artificial Analysis published results for two models from the same lab, Thinking Machines Lab: Inkling and Inkling Small. The smaller one beat the larger one on the real-work test. That same smaller model scores negative on a knowledge test — meaning it gets more answers wrong than right.

Both things are true at once. And understanding why is the most useful lesson a company can take away before choosing any artificial intelligence.

The Two Tests, in Plain Language

AA-Omniscience measures knowledge and hallucination. It is 6,000 questions across 42 economically relevant topics in six domains. The score runs from −100 to 100 and works like this: a correct answer adds, a wrong answer subtracts, and saying "I don't know" costs nothing. Zero means as many right as wrong.

AA-Briefcase measures work. The model receives realistic tasks across thousands of input files and has to deliver human things: spreadsheets, presentations, UI mock-ups. The evaluation looks at three dimensions — whether it is objectively correct, the quality of the analysis, and the quality of the presentation.

The Number That Should Worry You: 20%

On AA-Briefcase, Inkling Small scored 917 Elo against Inkling's 839. A comfortable win for the smaller model.

Except the objective check — the part that verifies whether the deliverable is factually right — was practically a tie: 20% versus 19%.

Read that again. In both cases, roughly one fifth of the objective checks passed. These are models that make headlines, on an office-work test, getting one in five checked items right. The gap between "winner" and "loser" did not come from being correct.

Why the Smaller One Won

It came from presentation. The report is explicit: Inkling Small's advantage was driven significantly by presenting the result better.

That is: the model that delivers a better-formatted spreadsheet, better-organised text and a prettier slide beats the model that gets the same number of facts right. Not because the evaluator is foolish — presentation is a legitimate dimension and it counts in the real world — but because, when correctness ties, looks decide.

Which is exactly what happens in your company's meeting room during an AI demo.

A Negative Score: More Wrong Than Right

On AA-Omniscience, Inkling landed at 2. Inkling Small landed at −9. A negative score literally means more wrong answers than right ones. Accuracy was 40% versus 31% — the smaller model knows less, which is expected with fewer parameters.

But look at the test's design, which is the clever part: there is no penalty for refusing to answer. A model that replies "I don't have that information" loses nothing. What drags the score down is answering wrongly with confidence.

So the scoreboard does not measure stupidity. It measures overconfidence — and that is the flaw that breaks an operation, because an admitted error gets handled, while an error delivered with assurance sails straight through.

What This Changes When Choosing AI for Your Company

The trap is the demo. Every AI tool demos well: whoever demos picks the example, and the chosen example always looks good. You watch, the fluency impresses you, and without noticing you are evaluating presentation — the very dimension that settled a benchmark where both contenders were right 20% of the time.

Three practical consequences:

  • Fluency is not competence. Well-written text and a correct answer are independent things. The model is trained to sound good, not to be right.
  • A bigger model does not always deliver better. It depends on the task: Inkling Small loses on knowledge and wins on work. Choosing by size is choosing by headline.
  • "I don't know" is a feature, not a weakness. In a business operation, the model that admits its limit is worth more than the eloquent one that invents.

How to Test Properly Without Becoming a Lab

You do not need to build a benchmark. You need a method, and it fits in an afternoon:

  1. Gather 20 to 30 real cases from your business — customer questions, documents, requests that actually happened.
  2. Write the correct answer before testing. Without an answer key, you will judge on fluency, like everyone else.
  3. Run the same cases through every candidate tool. Same cases, same order.
  4. Count three things separately: how many right, how many wrong with confidence, and how many times it admitted not knowing. The second column is the one that matters.
  5. Only then look at presentation. It is a tie-breaker, not a selection criterion.

One afternoon of testing with an answer key beats any published benchmark comparison — because it measures your case, not somebody else's.

Frequently Asked Questions

So benchmarks are useless?

They are very useful — as long as you read what they measure. The mistake is using a general score to decide a specific case. The two tests cited here are useful precisely because they separate correctness from presentation.

Is a small model always worse?

No. It tends to know fewer things, but it can execute a well-defined task better — and costs far less. For a good share of business uses, that is the right trade.

How do I reduce the risk of confidently wrong answers?

By giving the model context instead of relying on its memory: connecting it to your documents and data cuts invention sharply. And by stating explicitly, in the instruction, that it should say when it does not know.

What is the most common mistake when buying AI?

Choosing based on the demo. The demo shows the vendor's best case; your day-to-day is made of the worst ones.

Sources

Want help testing and choosing the right AI for your case?

More information