The maintenance technician stands at the line. The filler FL-200 shows error E-217. He asks the assistant: “Has this happened before, and what helped?” The answer arrives fast and sounds sure of itself. It explains in general terms how error codes arise on filling lines, points to a chapter in the manual, and stops there. The service report from spring, which names exactly this error and its fix (“inlet sensor readjusted”), never shows up. It is in the index. The search simply did not find it.
The example is invented, like every document in this article. The pattern behind it is not. The language model did nothing wrong here: it built a decent answer from the passages it was handed. The passages were wrong. This is where most of the disappointment with enterprise AI originates in practice, and it is where hardly anyone looks, because the failure is invisible. A bad answer looks like a bad answer. The fact that the right passage never reached the model does not show on the page.
Our thesis in this article: the quality of an enterprise search is decided by five choices made before the language model ever sees a word. Each one is easy to understand, each one can fail, and each one can be checked without a computer science degree. If you know them, you can question vendors better and prepare your own data better.
Why Search Matters More Than the Model
A language model does not know your documents. It does not know your framework agreement with the transport partner, the acceptance report of the filling line, or the service history of the filler. To let it answer from your knowledge anyway, you put a search in front of it. The search fetches the relevant passages and hands them to the model together with the question. The model writes the answer from them. In technical language this approach is called retrieval-augmented generation, or RAG.
A simple piece of arithmetic follows: the model can only answer as well as the passages it receives. If the decisive passage is missing, there are two bad outcomes. In the better case the model says it found nothing. In the worse case it fills the gap with something plausible. Both have the same cause, and neither is fixed by a larger model.
Thought experiment: you hand a very good lawyer a case file from which the most important page has been removed. He will work carefully and still reach the wrong conclusion. His care was not the problem.
The Chain in Five Stations
From file to answer, a document passes through five stations. The graphic lays them out side by side, with the two layers that run across them below: the index, which holds content and access list together, and the permission filter.
In detail:
- Chunking: The document is cut into passages, because you never want to search or hand over a whole document, only the one place inside it.
- Embedding: Each passage is converted into a sequence of numbers that represents its meaning. Similar meaning gives similar numbers.
- Searching: A keyword search and a meaning search each deliver a list of candidates, and the two lists are merged into one.
- Reranking: The best candidates are examined more closely and sorted again. Only the first few go to the model.
- Answering with sources: The model writes the answer and names the passages it came from, so that anyone can verify it.
Across all of this, one rule holds: what a person may not read must not influence that person's answer. More on that shortly. First, the stations where most searches stumble.
Station 1: Chunking Documents
Why split at all? For three practical reasons. A model can read only a limited amount of text at once. A single meaning value for a whole document would be so coarse that it distinguishes nothing. And the evidence under the answer should point to one place, not to 40 pages.
So the question is not whether to cut but where. Chunks that are too small lose their context. Chunks that are too large dilute it.
| Chunk Size | Typical Symptom | Example (fictional) |
|---|---|---|
| Too small (a single sentence) | The number is found, but nobody knows what it refers to | “34 minutes” sits alone in one chunk, the heading “format changeover” in another |
| Just right (one coherent thought) | The chunk reads on its own and fits a question | A paragraph on the format changeover with the agreed value and the measurement |
| Too large (a whole section or chapter) | The meaning is an average over many topics and the answer drowns | A 40-page chapter in which the format changeover takes up one paragraph |
Structure Beats Fixed Length
The obvious approach is to cut every document after a fixed number of characters. It is simple and works acceptably on uniform running text. On company documents it fails quickly, because they are not uniform. They have headings, tables, footnotes and appendices. A fixed length cuts through the middle of a table or separates a heading from the first paragraph below it.
It is better to cut along the structure: at headings and paragraphs, and for tables row by row, but with the header row repeated in every chunk. An example from the fictional acceptance report for project P-2291: the row “nominal output at 500 ml, 10,000 bottles per hour agreed, 10,240 measured” only makes sense with its column headers. Without them it is one number next to another. Add a small overlap between neighboring chunks, so that a thought that crosses a cut appears complete in at least one chunk.
You cannot see from the outside whether this worked. You see it when you have the chunks printed and read a few of them as if you were the search: do you understand each chunk without knowing the document? If not, the search will not understand it either.
Station 2: Embedding, or Meaning as Coordinates
An embedding model reads a passage of text and returns a long sequence of numbers, a vector. You can picture it as a coordinate on a map of meanings. Passages that say similar things lie close together, passages about other topics lie far apart. The question is converted into a coordinate in the same way, and the search returns the passages nearest to it.
The big advantage: it does not depend on the same words. Someone asking about the “notice period” also finds the paragraph that speaks of the “time limit for ending the contract”. A pure keyword search would miss it.
The weakness is the flip side: meaning is fuzzy. The model recognizes that “error E-217 on filler FL-200” belongs in the neighborhood of “fault on a filling line”, but the exact identifier E-217 is, for the map, just one digit sequence among many. Identifiers, article numbers, rare proper names and figures are precisely what meaning search holds on to worst. In the opening example, this was the first half of the failure.
Missing Context: The Passage Without a Sender
The second half sat in the passage itself. Picture the sentence “The period is 30 days.” in the fictional framework agreement with transport partner TP-118. Stored as its own chunk, it is worthless. Which period? In which contract? The chunk does not know, and the map cannot know either. It sits somewhere near “time limits”, together with thousands of sentences from other contracts.
The remedy is unspectacular and effective: before embedding, every chunk gets a context line with the document title, the section and, where it makes sense, the date. The sentence becomes “Framework agreement TP-118, termination: The period is 30 days.” Now it can be assigned unambiguously, and the same line can appear later in the source reference.
A word on data protection, because the question comes up regularly: embeddings are not automatically anonymous. Treat vectors the way you treat the documents they were made from. If the embedding runs at an external provider, that is usually commissioned data processing under Art. 28 GDPR, which has to be regulated by contract. With TheroAI, data is stored in Germany and AI processing takes place in the EU. This is not legal advice, so clarify your individual case with your legal counsel.
Station 3: Combining Keyword and Meaning
Keyword search is older and unfashionable, and it has a strength that meaning search lacks: it takes words literally. If you type “E-217”, you get places where exactly “E-217” appears. A common method for this is called BM25. It scores how often a search term occurs in a passage and how rare it is across the whole collection. Rare terms such as an error number count for more than common ones such as “error”.
Meaning search has the opposite strength. It finds paraphrases, synonyms and phrasing from another department. Both methods therefore have gaps, just in different places. That is why modern enterprise searches combine them. This is called hybrid search. The search in TheroAI works this way too: keyword and meaning run side by side.
Abbreviations and Spellings
Abbreviations in particular are the stumbling block in everyday operations. Every company has its own shorthand that is explained nowhere. Sales writes “SLA”, the contract says “service level agreement”. A line is called “filling line” in the requirements specification and “filler” in the service report. One person writes the identifier “E-217”, the next “E217”, a third “E 217”. To the search, these are three different words.
Three things help, and they work together: normalizing spellings before indexing, a list of abbreviations and synonyms that you maintain yourself, and the context line from station 2, which often already contains the abbreviation in spelled-out form. None of the three is magic. All three are manual work that pays off.
Two Lists Become One
Keyword search and meaning search each return a ranking with their own scores, and those scores cannot be compared directly. A common way out is not to mix the scores but the ranks. The method is called reciprocal rank fusion. Each list awards a rank a value of 1 divided by (60 plus rank), and the values from both lists are added up. The 60 is a common default, and it keeps rank 1 from dominating disproportionately.
What matters about the method: someone who is reasonably far up in both lists beats someone who is far up in only one. A worked example with three documents shows it.
| Document | Keyword Rank | Meaning Rank | Points (times 1,000) |
|---|---|---|---|
| B | 3 | 2 | 32.0 |
| A | 1 | 40 | 26.4 |
| C | 2 | not in the list | 16.1 |
Document A tops the keyword list but ends up behind B, because meaning search only sees it at rank 40. Document C appears in just one list and falls behind accordingly. The numbers are a worked example with assumptions, not a measurement.
Station 4: Reranking the Hits
After the fusion there is a list of candidates, roughly sorted. Roughly means: the methods before this compare question and passage only indirectly. The passages were converted into coordinates in advance, the question is converted only now, and the two meet only through their distance. That is fast, but imprecise.
A reranker works differently. It reads question and passage together, the way a person would lay the two side by side, and judges whether the passage really answers the question. That is considerably more precise and considerably more expensive, because it has to happen for each pair individually and cannot be computed in advance. So you apply it only to a small set: the best candidates from the fusion, not the whole collection.
The funnel shows the orders of magnitude in a worked example. The assumptions are in the open in the table.
| Stage | Example Value (assumption) | What Happens There |
|---|---|---|
| Index | 200,000 chunks | All ingested chunks from all sources |
| After permission filter | 60,000 chunks | Only chunks this person may read |
| After search and fusion | 50 candidates | Keyword and meaning, merged |
| After reranking | 5 chunks | Question and chunk judged together |
The reranker proves its value in the cases where the fusion delivers two almost equally good candidates. At rank 3 sits, for example, the manual chapter that explains the error code only in general terms, at rank 9 the service report with the concrete fix. A reranker that reads the question “Has this happened before, and what helped?” together with both passages recognizes that only the report answers it, and moves it forward. Without this step, the service report might never have reached the model, because only the first few ranks are passed on.
Permissions: The Filter Belongs Before the Ranking
The fifth choice is not a station but a rule for all of them: a person may only be answered from passages that person may read. What matters is when the filter applies. If it is applied only to the finished hit list, forbidden documents take up slots that a permitted document would have needed, and excerpts can be created before the filter runs. If it applies inside the search, the ranking is computed over permitted passages only, from the very start. The funnel above shows it: the permission filter stands right at the front, before the first ranking.
For this, the access lists of the source have to reach the index together with the content and be kept current. In the setup of the Google Drive connector in TheroAI this is visible: indexing user permissions is its own ticked item there, next to updating content and permissions several times per hour. You can read more on the page about connectors.
There is an honest limit here too: the index is only as current as the last synchronization. When a share is revoked in the source, there is a time window before it is mirrored in the index. How long it is depends on the source and on the sync interval. Ask every vendor for this value, because the answer “synchronized” alone is not enough. How TheroAI mirrors permissions in detail is described in our article on permission mirroring. For data storage and security as a whole, see security.
The Evidence: Answer with Sources
At the end of the chain stands the answer, and with it the question of whether you can check it. A search that has found well can show it: every statement carries a number, and the number leads to the place it came from.
The fictional example from plant engineering shows what that looks like. Asked what was agreed at the acceptance of the filling line in project P-2291 and what was not met, the assistant answers with a table of criterion, agreed value, measured value and result. Three criteria are met, the format changeover is not: at most 30 minutes were agreed, 34 were measured. On the right, the source panel opens with the acceptance report, document AP-2291-02, dated 17.10.2025.
That number is the return ticket through the whole chain. If the answer is wrong, you can trace which passage the model received, whether it was cut correctly, whether the context line was right, and whether it stood at a sensible rank. Without evidence, all that is left is a gut feeling. Industry examples for this can be found on the page for plant engineering.
Where Searches Fail: Seven Patterns
In practice, the same failure patterns keep recurring. The overview below assigns them to the stations and names how to recognize them in daily work.
| Pattern | Station | How You Notice | Remedy |
|---|---|---|---|
| Chunks too large | Chunking | Answers are general, the concrete passage is missing | Cut along structure, smaller topical chunks |
| Chunks too small | Chunking | Numbers without context, the answer contradicts itself | Bundle sentences that belong together, slight overlap |
| Missing context | Embedding | Identical sentences from different documents get confused | Context line with title, section and date |
| Identifiers get lost | Searching | The search returns topics, but not the exact number | Enable keyword search, unify spellings |
| Unknown abbreviations | Searching | Whoever uses the shorthand finds nothing, whoever spells it out does | A maintained abbreviation and synonym list |
| Wrong order | Reranking | The right document is there, but at rank 9 | Reranker on the best candidates |
| Filter too late | Permissions | Answers do not change after a share is revoked, or hits go missing for no reason | Filter inside the search, short sync intervals |
Two patterns deserve an additional note. The first is outdated versions: if three versions of a contract are in the collection, the search finds all three and has no reason to prefer the newest. The date in the context line helps here and, where binding force matters, a clear place for the valid version. The second is tables and scans. A table filed as an image contains no text for a text search until recognition extracts it. Ask a vendor how tables and scanned documents are handled.
How to Test Your Own Search
You cannot judge a search from a demo alone, because demos run on questions the vendor knows well. An honest test takes an afternoon and needs no software.
- 1.Collect 20 real questions from daily work, not 20 nice ones. Mix questions with identifiers, with paraphrases and with abbreviations.
- 2.Note in advance for each question which document and which passage holds the answer.
- 3.Have the questions asked and check not the answer first but the sources: is the right passage among the first five?
- 4.Count separately by question type. If all identifier questions fail, that points to missing keyword search. If the paraphrases fail, meaning search is missing.
- 5.Repeat the test after every major change to filing or settings. A search does not become good once. It stays good or gets worse.
The test measures nothing universal. It tells you how a search copes with your documents and your questions, and that is the only number that counts for you. It also shows an honest result nobody likes to hear: no search finds everything. What matters is what it does when it finds nothing. A good assistant says so instead of inventing a plausible answer.
What You Can Take Away
If you are evaluating or building an enterprise search, these check questions are worth asking:
- How are documents split, by fixed length or along their structure, and do table headers stay with the rows?
- Does every chunk get a context line with title, section and date before it is embedded?
- Does a keyword search run next to the meaning search, so that identifiers and abbreviations are found?
- Are the candidates reranked before the model, and how many chunks does the model receive in the end?
- Does the permission filter apply before the ranking, and how quickly does a revoked share arrive in the index?
- Does every statement show a source that you can open and compare with the passage in the document?
Our position: a good enterprise search is not one clever idea but a chain of ordinary decisions, none of which shines on its own and each of which can fail individually. If you know the chain, you do not rely on the claim that “the AI finds everything” but check where it holds. That is less spectacular than a miracle model and considerably more reliable.
If you would like to see how this looks with your own questions, we are happy to show you in a short demo, without pressure and with your examples.
See Thero live
Book a short demo. You talk directly to the founding team.