洞见
约 3 分钟阅读作者 AiHPC
"SOTA on the leaderboard" vs "works on our work"
"SOTA on the leaderboard" vs "works on our work"
TL;DR. SOTA on a public leaderboard means a model did well on someone else's test. It does not mean it will follow your SOPs, read your Cantonese notes, or stay inside your compliance walls. A short, honest eval on your work — including publishing a weak result — beats a shiny demo slide.
What a leaderboard is good for
Public benchmarks are useful the way a racing circuit is useful:
- They let researchers compare systems under a shared rulebook.
- They show rough capacity trends over time.
- They give journalists and investors a number to quote.
That is real value. It is still a shared exam, not your job interview.
What a leaderboard cannot tell you
Your organisation's work usually fails in ways the public set never saw:
| Your reality | Why the scoreboard misses it |
|---|---|
| Internal SOPs and forms | Not in the public test set |
| Mixed English + 繁體 / Cantonese notes | Many suites are English-heavy |
| Permissioned documents | Leaderboards do not encode your RBAC |
| Multi-step agent tools | Trajectory failures ≠ single-answer quizzes |
| "Refuse when unsure" | Some suites reward confident guessing |
So a vendor can be state-of-the-art on the board and still fail the first week on your helpdesk, your clinical form, or your procurement checklist.
The demo slide problem
Demos are optimised to look good:
- Cherry-picked prompts.
- Yesterday's tidy corpus.
- A human quietly fixing the awkward turn.
- No published failures.
Buyers who ask "will this work here?" need the opposite habit: a small golden set of real tasks, scored the same way every time, with weak results kept in the open. That habit is how you stop buying theatre.
A calm checklist for "works on our work"
Bring this to the next vendor meeting (or your own pilot review):
| Clarify | Good enough answer sounds like… |
|---|---|
| Tasks | "These 20–50 real questions / workflows are in scope for v1." |
| Pass rule | "Pass means X — citation, tool call, refusal — not 'sounds fluent'." |
| Languages | "We scored the languages you actually write in." |
| Hosting | "The eval ran in the same residency posture as production." |
| Regression | "When we change the model or prompt, we re-run the same suite." |
| Honesty | "Here are the failures — we are not hiding the weak week." |
Notice what is not enough: a higher public rank, a bigger "B", or a longer feature PDF.
How this connects to what we build
OrchAI Eval is our name for the quality inspector in the suite: eval suites that can hold Library, Agents, and AI Apps to a line before a release reaches users — the CI/CD gate framing, not "trust the demo." The literacy point does not wait for every dashboard to ship: judge AI on your work. Soften any vendor claim that only cites a public board.
If you are past "SOTA on the slide" and ready to write a short golden set, talk to us or try the demo.
Frequently asked questions
Should we ignore leaderboards? No — use them as a rough capacity signal. Then eval on your tasks and failure modes.
What is a short honest eval? A small golden set, clear pass/fail, same run every time — and weak results published. Hidden failures are marketing.
How does OrchAI Eval fit? The suite's quality inspector — suites that can block regressions before users see them. Judge on your work, not only on a public scoreboard.
常见问题
此页面尚未提供你的语言版本,现以英文显示。