洞见

约 3 分钟阅读作者 AiHPC

"SOTA on the leaderboard" vs "works on our work"

ai-literacy
kepu
eval
buying-ai
benchmarks

"SOTA on the leaderboard" vs "works on our work"

TL;DR. SOTA on a public leaderboard means a model did well on someone else's test. It does not mean it will follow your SOPs, read your Cantonese notes, or stay inside your compliance walls. A short, honest eval on your work — including publishing a weak result — beats a shiny demo slide.

What a leaderboard is good for

Public benchmarks are useful the way a racing circuit is useful:

  • They let researchers compare systems under a shared rulebook.
  • They show rough capacity trends over time.
  • They give journalists and investors a number to quote.

That is real value. It is still a shared exam, not your job interview.

What a leaderboard cannot tell you

Your organisation's work usually fails in ways the public set never saw:

Your realityWhy the scoreboard misses it
Internal SOPs and formsNot in the public test set
Mixed English + 繁體 / Cantonese notesMany suites are English-heavy
Permissioned documentsLeaderboards do not encode your RBAC
Multi-step agent toolsTrajectory failures ≠ single-answer quizzes
"Refuse when unsure"Some suites reward confident guessing

So a vendor can be state-of-the-art on the board and still fail the first week on your helpdesk, your clinical form, or your procurement checklist.

The demo slide problem

Demos are optimised to look good:

  • Cherry-picked prompts.
  • Yesterday's tidy corpus.
  • A human quietly fixing the awkward turn.
  • No published failures.

Buyers who ask "will this work here?" need the opposite habit: a small golden set of real tasks, scored the same way every time, with weak results kept in the open. That habit is how you stop buying theatre.

A calm checklist for "works on our work"

Bring this to the next vendor meeting (or your own pilot review):

ClarifyGood enough answer sounds like…
Tasks"These 20–50 real questions / workflows are in scope for v1."
Pass rule"Pass means X — citation, tool call, refusal — not 'sounds fluent'."
Languages"We scored the languages you actually write in."
Hosting"The eval ran in the same residency posture as production."
Regression"When we change the model or prompt, we re-run the same suite."
Honesty"Here are the failures — we are not hiding the weak week."

Notice what is not enough: a higher public rank, a bigger "B", or a longer feature PDF.

How this connects to what we build

OrchAI Eval is our name for the quality inspector in the suite: eval suites that can hold Library, Agents, and AI Apps to a line before a release reaches users — the CI/CD gate framing, not "trust the demo." The literacy point does not wait for every dashboard to ship: judge AI on your work. Soften any vendor claim that only cites a public board.

If you are past "SOTA on the slide" and ready to write a short golden set, talk to us or try the demo.

Frequently asked questions

Should we ignore leaderboards? No — use them as a rough capacity signal. Then eval on your tasks and failure modes.

What is a short honest eval? A small golden set, clear pass/fail, same run every time — and weak results published. Hidden failures are marketing.

How does OrchAI Eval fit? The suite's quality inspector — suites that can block regressions before users see them. Judge on your work, not only on a public scoreboard.

常见问题

与我们联系

告诉我们你的场景——医院、政府或金融。

联系 AiHPC

认识 OrchAI

为受监管买家而设的治理型 AI 平台。

探索 OrchAI

其他语言版本 en · 繁体中文

此页面尚未提供你的语言版本,现以英文显示。

AiHPC Innovation Limited · 香港 + 台湾 + 新加坡