Chapter Zero · Beginner FAQ

Why Does the "#1 on the Leaderboard" Model Feel Worse in Real Use?

You see a headline that a new model "dominates the leaderboard," switch over in a rush — and writing a weekly report still feels less smooth than before. You didn't use it wrong, and the board isn't fake. The test and what you need were never aligned.

One-sentence answer

The leaderboard tests an exam; you need it to do work. The question bank can be gamed, and the syllabus may not include your scenario — treat scores as a reference. The most reliable way to pick a model is to try it on your own work.

Reversal Demo · Who Beat the High Scorer

Here are two fictional models: A scores 95, B scores 88. By the leaderboard you'd pick A. But hang on — tap the three real tasks below and see how each one actually does.

Model A

95 pts
High on the leaderboard, a regular in the news

Model B

88 pts
Middling score, nobody writes about it
You've tried all three tasks. The conclusion is hard to miss: the gap between 95 and 88 may be invisible in your everyday work — or even reversed. Scores rank seats in the exam hall. How well it works still has to be tried by hand.
Why This Happens · Three Reasons
📖

Leaderboard gaming

Once a leaderboard's question bank has been floating around the internet long enough, it inevitably leaks into training data. It's like memorizing the answers before the exam: the score looks great, the ability isn't there. The industry has a polite name for this: "data contamination."

🎯

Overfitting the exam

Vendors know everyone watches the board, so they optimize specifically for the test points. Like an exam-prep grind student: top of the class on paper, while everyday conversation and hands-on skill may get sacrificed.

🗺️

The syllabus doesn't include your scenario

Leaderboards test math, code, and knowledge Q&A — but not "does it get your industry", and not "can it comfort someone." The subject you care about most may not be on the syllabus at all.

Which Boards Are Relatively Trustworthy · Look for Human Blind Tests

Some boards are relatively trustworthy: human blind-test voting. Two models' answers are shown side by side, names hidden, and a large number of real users vote for the better one. There's no fixed question bank to memorize, the judges are living people, and gaming it is much harder.

But it's only "relatively" trustworthy: the voters may not work in your field, and a style the public likes may not fit your scenario. To see how Chinese models actually perform in real scenarios, head over to How to Choose a Chinese Model.

How to Choose for Yourself · Build a Private Question Bank

The most reliable evaluation lab is you. The method is unglamorous, but it works extremely well: collect 10 real questions from your own work and save them as a document. Every time you want to switch models, run the 10 questions. Whether it's usable is obvious at a glance.

How to build a private question bank (examples)

This question bank has a hidden bonus: it forces you to get clear on what you actually need AI for. Plenty of people switch models for three months, then realize the real problem was a poorly written prompt. With the bank in hand, switching models takes five minutes to decide — no more chasing the news.

✅ What this page wants to share with you