The answer exists. Finding it is the job.
Contracts, specifications, invoices and reports hold what somebody needs to know, and getting it out means a person opening files until they find the right one.
What it looks like from the inside
Most organisations do not have a knowledge problem. They have a retrieval problem. The clause is in the contract, the figure is in the report, the exception is in an email from March. Somebody has to know which document to open, and that somebody is usually one person who is busy.
Keyword search fails here because the document says “termination for convenience” and the person asked “can we get out early”. Vector search alone fails the other way: it returns something close in meaning and nothing exact, including a figure that is nearly right. What holds is both, with a citation, so the reader checks the answer against the page rather than trusting it.
The harder half is the documents themselves. A scan is not text. A table read as prose loses the column it belonged to. A quantity extracted without its unit is worse than no answer at all, because it looks usable.
- Finding the clause takes longer than acting on it
- Two people read the same document and quote different numbers
- The answer arrives by asking a colleague rather than by searching
- Scanned files are effectively invisible to every search you have
- Data is retyped from a PDF into a system that could have received it
What we do about it
Work out what is actually asked of the documents
We start from the questions people bring, not from the file count. The shape of those questions decides whether this is retrieval, extraction, or a form somebody should have been given years ago.
Get the text out properly
Layout, tables and scans handled as structure rather than as a wall of characters, so a figure keeps the column and the unit it arrived with.
Retrieve with both methods, and cite
Hybrid keyword and vector search with reranking, tuned on your own questions. Every answer carries the document and the passage it came from.
Score it before anyone depends on it
An evaluation set built from real questions with known answers, run before launch and on every change. Where the system should say it does not know, that is a correct answer and it is scored as one.
What you end up with: search and extraction over your own documents, a citation on every answer, an evaluation set drawn from real questions, and the score it produces
The services this draws on
Other weeks that are not working
Does this sound like you?
Send a paragraph about how it shows up in your week. We reply within two working days with questions, a rough shape, and an honest answer on whether we are the right people for it.
Discuss a problem