洞见
约 3 分钟阅读作者 AiHPC
Do 1M context windows equal perfect memory?
Do 1M context windows equal perfect memory?
TL;DR. No. A 1M context window is a bigger place to paste for this sitting. What teams actually want is: ask a plain question over their files, see which page the answer came from, and keep those documents inside the building. That is OrchAI Library — not a million-token chat box.
What buyers say they need
The conversation almost always sounds like this:
I have thousands of documents — reports, guidelines, contracts, papers — and the answers I need every day are buried in them. I want to ask in ordinary language and get a straight answer. I cannot paste confidential files into a public chatbot, and I do not have time to build this myself.
That sentence is the job. A bigger window does not finish it.
Why a 1M paste feels like the shortcut
A long context window is useful when the pack is already in your hand: one fat contract, one transcript, two versions you chose on purpose. You paste once. The model sees more of this sitting. Fine.
The shortcut people hope for is: dump the department's files into a 1M chat and call it memory. That is still a pile on the desk — not a librarian who has read the shelves and can point to the page.
Three reasons the shortcut breaks:
- The middle of a long paste is easy to miss. Models, like tired readers, use the start and the end more reliably than the stack in between. Public research has a name for this (lost in the middle). It is not an AiHPC score. It is why "it fitted in the window" is not "it was used."
- Monday's paste is not Tuesday's procedure. The window holds a snapshot. Your SOPs move. The chat does not notice unless someone pastes again.
- A fluent paragraph is not a source you can open. Buyers who must defend an answer need the document and the page — click, read, check. And for a hospital, a government desk, a law firm, or a lab, those files often must not leave the building.
What we actually sell
OrchAI Library is the memory of the suite. The line we use with buyers:
Private Google for your team's documents — every answer cited to the source.
You point it at your folders. People ask in plain language. Every answer shows which document and which page it came from. Click the citation. Read the source. No guessing whether the model made it up.
Why that is the purchase, not a bigger paste box:
- It cites. Grounded in a real file you can open. Not hallucination roulette.
- It stays put. Documents on your premises. Air-gap-friendly. Nothing goes to a public cloud unless you choose that.
- It works in a day. A small server — even a Mac mini — not a platform programme. Library is the small, fast tier; Portal is the whole building when you need enterprise governance.
Chat is a surface. The product is the cited-answer engine underneath. Library finds and cites. If you want software that acts, that is Agents. If you want the quality habit in front of a change, that is Eval (still being built — monitoring and a pre-ship gate, not a finished customer harness).
Honest limit: asking and citing on Library is the path we demo. Document upload can still sit behind Portal chrome — we do not sell "upload only in Library and never touch anything else" as a finished happy path.
The offer, in one sitting
If your team is 5–50 people drowning in their own PDFs — a lab, a specialty clinic, a boutique firm — you do not need a million-token slogan. You need answers you can check, on hardware you control.
See also
常见问题
此页面尚未提供你的语言版本,现以英文显示。