Why Does the "#1 on the Leaderboard" Model Feel Worse in Real Use?
You see a headline that a new model "dominates the leaderboard," switch over in a rush — and writing a weekly report still feels less smooth than before. You didn't use it wrong, and the board isn't fake. The test and what you need were never aligned.
The leaderboard tests an exam; you need it to do work. The question bank can be gamed, and the syllabus may not include your scenario — treat scores as a reference. The most reliable way to pick a model is to try it on your own work.
Here are two fictional models: A scores 95, B scores 88. By the leaderboard you'd pick A. But hang on — tap the three real tasks below and see how each one actually does.
Model A
95 ptsModel B
88 ptsLeaderboard gaming
Once a leaderboard's question bank has been floating around the internet long enough, it inevitably leaks into training data. It's like memorizing the answers before the exam: the score looks great, the ability isn't there. The industry has a polite name for this: "data contamination."
Overfitting the exam
Vendors know everyone watches the board, so they optimize specifically for the test points. Like an exam-prep grind student: top of the class on paper, while everyday conversation and hands-on skill may get sacrificed.
The syllabus doesn't include your scenario
Leaderboards test math, code, and knowledge Q&A — but not "does it get your industry", and not "can it comfort someone." The subject you care about most may not be on the syllabus at all.
Some boards are relatively trustworthy: human blind-test voting. Two models' answers are shown side by side, names hidden, and a large number of real users vote for the better one. There's no fixed question bank to memorize, the judges are living people, and gaming it is much harder.
But it's only "relatively" trustworthy: the voters may not work in your field, and a style the public likes may not fit your scenario. To see how Chinese models actually perform in real scenarios, head over to How to Choose a Chinese Model.
The most reliable evaluation lab is you. The method is unglamorous, but it works extremely well: collect 10 real questions from your own work and save them as a document. Every time you want to switch models, run the 10 questions. Whether it's usable is obvious at a glance.
How to build a private question bank (examples)
- Pick the work you do most: weekly reports, proposals, emails — choose 3 things you do every week
- Pick the most specialist work: questions with industry jargon and internal context, to see if it gets your field
- Pick the work that went wrong before: questions AI already botched — the best test of a new model's quality
- Keep one easy giveaway and one nasty question: if it fails the giveaway, drop it; the nasty one is for separating the pack
- You decide what a good answer looks like: you know better than any leaderboard what output counts as "good enough to submit"
✅ What this page wants to share with you
- Benchmark scores are the entrance exam; you need the probation-period performance: the two are related, but only so much
- A high score may mean it memorized the questions: the bank leaks into training data, and test points get targeted optimization
- Human blind-test voting boards are relatively trustworthy: no fixed question bank, the judges are living people
- Build your own 10-question quiz: run it when you switch models, get an answer in five minutes — more useful than chasing the news