It is Monday, a little before nine. In the sales team of a plant manufacturer, a customer asks whether the acceptance test of the filling line in project P-2291 went through without complaints. A colleague puts the question to the assistant. The answer arrives within seconds, in clean prose and in the right tone: “All four acceptance criteria were met, the line has been accepted without reservation.” He pastes the sentence into his email and hits send.
The sentence is wrong. The acceptance report says that the format changeover took 34 minutes instead of the agreed maximum of 30, and that the line was therefore accepted only with a reservation. The scene is made up, and the documents come from our fictional demo data. But every company that uses language models has a version of it.
The first reflex is to ask for a better model. This article argues for something else. A better model makes mistakes rarer, but not more visible. What works are three rules that apply around the model: the assistant names its sources (mandatory sources), every single statement carries its own citation, and when the documents do not contain the answer, the assistant says so plainly: “Not in the documents.” On top of that comes a review routine for you as the reader. It covers four kinds of error and costs seconds per statement.
Why a Better Model Is Not Enough
“Hallucination” is a convenient word, but an imprecise one. What it describes is a statement that reads fluently and rests on nothing real. Language models produce text that fits the text before it. When a piece of information is missing, a plausible sentence is often the most natural continuation, while “I don't know” often is not. Better models and better training can reduce this. In our view it will not disappear entirely, and that is not even the decisive point.
The decisive question is a different one: what happens when the error occurs anyway? A wrong sentence looks like a right one. You read it with the same impression of certainty. The better models get at phrasing, the less the text itself reveals whether it is true. That is a line of reasoning, not a measurement, but it has an uncomfortable consequence: a better model can even make the situation worse, because the remaining errors are harder to notice and readers' trust grows faster than the reliability does.
There is a further point. Part of the typical errors have little to do with language ability and a lot to do with the documents and with how things are matched. A model cannot know that the 2023 price list was replaced by a newer one if both sit in the archive and nothing marks the difference. All it can do is show where it got something from. Then you decide.
| Question | A better model | Mandatory sources with one citation per statement |
|---|---|---|
| How often is the answer off? | Tendentially less often | Does not change the frequency directly |
| Can you spot an error while reading? | No, wrong sentences sound just as sure | Yes, a missing or mismatched passage stands out |
| What happens with outdated documents? | They stay outdated and nobody sees it | The date of the source is visible |
| What happens at a knowledge gap? | The model may keep trying to fill it | “Not in the documents” is an allowed answer |
| How much work is the check? | Unchanged | Drops, because the path to the passage is given |
This leads to the thread running through the article. We cannot bring the error rate to zero, but we can lower the price of finding an error. Mandatory sources, one citation per statement and the honest gap all do exactly that. They turn a trust problem into a checking problem, and checking problems can be organized.
Four Ways an Assistant Gets It Wrong
If you want to catch errors in daily work, you need to know what you are looking for. In practice four patterns show up, and they behave differently when you check them. The examples below are invented and loosely based on our fictional demo documents. They are not customer cases.
Invented
The statement has no basis in the documents. The model added something that sounds plausible: a number, a clause, a commitment. In the opening it was the sentence “All four acceptance criteria were met.” The documents say the opposite. Invented statements are easiest to catch when every statement has to carry a citation: you spot them because the number is missing or because the source does not contain the statement.
Outdated
The statement used to be true. Suppose error E-217 on the filler FL-200 is explained with instructions from an older edition of the maintenance manual, while a revised edition prescribes a different procedure. The assistant quotes correctly, just the wrong edition. The model cannot resolve this by itself as long as both documents sit in the archive with equal standing. You spot it from the date of the source and from the fact that there are several versions.
Misattributed
The statement really is in the documents, but it belongs to a different subject. Example: asked about the offer for the filling line P-2417, the assistant gives a rated output of 10,000 bottles per hour. The value comes from the acceptance report of P-2291. Both projects live in the same archive, and both documents are about filling lines. Mix-ups like this usually involve identifiers, names and figures: P-2291 instead of P-2417, E-217 instead of E-271, one customer instead of another. You catch them by comparing the identifier in the answer with the identifier in the document.
From Another Context
The statement sits in the right document and belongs to the right subject, but it applies under different conditions. Suppose the framework agreement with the carrier TP-118 contains a delivery period that explicitly applies to standard goods only. The assistant passes it on as a general period, and a customer with a custom build gets the wrong information. Drafts, examples, exceptions and deadlines with conditions belong here too. This is the sneakiest error, because citation, date and subject all check out. You only catch it by reading the passage in context: the heading, the conditions, and words like “only”, “provided that” or “notwithstanding”.
| Kind of Error | Example (invented) | How You Recognize It | Review Question |
|---|---|---|---|
| Invented | “All four acceptance criteria were met” although the format changeover was not | The number is missing, or the source does not contain the statement | Does the passage exist? |
| Outdated | Fix for error E-217 taken from an older edition of the manual | The date or edition of the source is old, a newer one exists | Is it current? |
| Misattributed | A value from P-2291 is given for P-2417 | The identifier in the answer differs from the identifier in the document | Does it belong to the right subject? |
| From another context | A delivery period from TP-118 that applies to standard goods only is passed on as general | The condition or exception sits in the paragraph before or after | Does it apply to my case? |
The difference between the kinds explains why a plain list of sources at the end of an answer is not enough. In the first kind, the citation is missing. In the other three there is a citation, it just points to the wrong version, the wrong subject or the wrong conditions. Anyone who only checks whether some source is named will miss exactly those three.
Mandatory Sources Means Every Statement Carries Its Citation
Mandatory sources sounds like a footnote at the end. What is meant is stricter. How to recognize a usable source obligation:
- One citation per statement, not a list at the end. A list of five documents does not tell you which sentence comes from where.
- The source is a real document from the archive that you can open, not a reference the model wrote freely.
- The source shows what you need for the check: title, document number and date.
- A statement without a citation is recognizable as such and counts as unchecked.
- The search is limited to documents the asking person is allowed to see.
- The system warns when the quoted passage supports the statement only weakly.
In TheroAI it looks like this: the answer about the acceptance of P-2291 contains a table with the columns Kriterium (criterion), Vereinbart (agreed), Gemessen (measured) and Ergebnis (result). Behind each row sits a superscript number. On the right a “Quellen” (sources) panel opens with the entry “Abnahmeprotokoll Abfüllanlage P-2291, Dokument AP-2291-02, Datum 17.10.2025” (acceptance report filling line P-2291, document AP-2291-02, dated 17 October 2025). The answer itself states the values: the format changeover took 34 minutes, the agreement was a maximum of 30, and the result reads “nicht erfüllt” (not met). The other three criteria, rated output, filling accuracy and availability, are marked “erfüllt” (met) in the same table. A click on a number takes you to the place in the document where the figure comes from.
Title, document number and date are exactly what you need for the questions “Is it current?” and “Does it belong to the right subject?”. You do not have to go looking, they sit next to the answer. TheroAI can also mark a citation with a hint when the quoted passage supports the statement only weakly. That does not replace your check. It only points you to where it is most urgent.
The Path of an Answer
A source obligation is credible when it is not bolted on at the end but shapes the whole path of the answer. The graphic shows in general terms what that path should look like, in four stations:
First, the question goes to a search, not straight to the memory of the model. Second, the search works only in documents the asking person has access to. Third, the answer is built from the passages found, and every statement gets the number of its passage. Fourth, the sources wait with document and date, ready to open. And for the case where the search finds nothing suitable, there is a separate exit, described in the next section. The rule we consider right: a statement without a passage does not get written.
The Honest Answer: “Not in the Documents”
The most effective sentence against hallucinations is one that many systems are reluctant to say. A language model is built to sound helpful, and in text a gap is not a natural place to stop. That is why a rule is needed: if the documents do not answer the question, the assistant says so. It does not guess and it does not round up to “usually”.
A good gap message is more than “I don't know”. It says what was searched, names related material as related material and leaves the decision to you. The table shows four situations and the difference between a poor and a good answer.
| Situation | Poor Answer | Good Answer |
|---|---|---|
| Nothing found | A plausible value that sounds “usual” | “Nothing on this in the documents searched”, stating what was searched |
| Only related material found | The related material is rephrased as the answer | The related material is named as such: “I found X, which does not answer your question” |
| Two documents contradict each other | One version is silently picked | Both passages are named with their dates, and the decision stays with you |
| Only an older version exists | The old version counts as current | Version and date are named, with a note that it may be superseded |
The sentence only works if your company can bear it. If employees read a gap message as a failure of the tool, they will reach for the tool that always answers, and with it for the one that is wrong more often. So treat every gap message as information: it shows that something is not documented. Whoever maintains the documents should see these messages regularly.
Where May the Answer Come From?
Part of an honest answer is where it may come from. In TheroAI you choose this in the chat under “Quellen” (sources). There are five options: “Automatisch” (automatic), “Nur Unternehmenswissen” (company knowledge only), “Nur öffentliche Quellen” (public sources only), “Unternehmen und öffentliche Quellen” (company and public sources) and “Keine externen Quellen” (no external sources).
For questions about contracts, acceptance tests or faults, “company knowledge only” is our recommendation: what is not in there cannot quietly come from the web, and a gap message stays a gap message. According to the menu, “no external sources” means answering from the language model alone. Such an answer cannot carry a citation. Treat it as a draft until you have checked it against a source.
Permissions are part of the honest answer too. The assistant searches only what the asking person may see. “Not in the documents” therefore means: not in the documents you have access to. Whether a document exists somewhere else is something an assistant should not reveal without undermining access rights. More on this on the security page.
What a Check Costs: An Example Calculation
Checks that are too laborious do not happen. That is the real mechanism behind mandatory sources: they do not make the check unnecessary, they make it cheap enough that it actually takes place in daily work. An example calculation with round assumptions shows the order of magnitude. The values are assumed, not measured.
| Assumption | Without a Citation per Statement | With a Citation per Statement |
|---|---|---|
| Statements in one answer | 10 | 10 |
| Time per statement | 3 minutes (own search in the archive) | 20 seconds (open the citation and read the passage) |
| Time for the whole answer | 30 minutes | 3 minutes 20 seconds |
Ten statements at three minutes each make 30 minutes. Ten statements at 20 seconds each make 200 seconds, that is 3 minutes and 20 seconds. The ratio is 9 to 1. Your own values will differ, the direction will not: when the citation sits next to the statement, the search disappears, and the search is the biggest part of the work.
That does not mean you should check every statement of every answer. How deeply you check depends on what you use the answer for. That is the subject of the next section.
The Review Routine for Everyday Work
The four kinds of error produce four questions. Run through them for every statement you intend to pass on:
- 1.Does the passage exist? Open the citation and look for the statement there. If the citation or the passage is missing, treat the statement as invented until proven otherwise.
- 2.Is it current? Check the date and version of the source. Is there a newer edition, an addendum, an amendment?
- 3.Does it belong to the right subject? Compare the identifiers in the answer and in the document: project, customer, device, error code, contract.
- 4.Does it apply to my case? Read the paragraph before and after. Watch for conditions, exceptions and words like “only”, “provided that” and “notwithstanding”.
Numbers, dates, names, deadlines and commitments deserve special attention. That is where a mistake costs the most and is easiest to overlook. A sentence that summarizes the course of a conversation is less critical than a deadline.
Would I sign this statement if the assistant had not been involved?
This one question does not replace any of the four, but it helps you pick the depth of the check. Whoever signs something opens the source.
Review Depth by Consequence
Not every answer needs the same care. A simple rule of thumb follows the consequences of a mistake:
The top level can be secured organizationally instead of being left to the goodwill of each individual. A workflow can create the draft answer and then pause at an approval before it is released. The workflows in TheroAI have an approval step for this, with a deadline of 72 hours and a switch for the four-eyes principle. The draft of a customer reply names its sources at the end, so the reviewing person has them in front of them.
What the Company Has to Provide
Mandatory sources are not a setting you switch on and forget. Three things belong with it.
Documents that can be checked. The date of the source only helps if it is right. Put version and date in the title or properties of documents, mark replaced versions as replaced or file them away, and name one responsible person per collection. An “outdated” answer is often a maintenance problem in the archive and not a model problem.
A route for gaps. Every “not in the documents” message points to a documentation gap or to a search that is set too narrowly. Decide who receives these messages and what becomes of them.
Clear responsibility. Whoever passes statements from an assistant on to customers, authorities or employees remains responsible for the content. Which checking and documentation duties apply in a given case depends on the industry and the contract. This is not legal advice, so clarify the individual case with your legal counsel. If you log reviews, that may touch the behavior or performance of individual employees. In Germany the works council may then have a say (Betriebsverfassungsgesetz section 87 paragraph 1 number 6). You clarify that case by case as well.
How You Know It Works
Whether the rules take hold shows not in a feeling but in a few figures that you can collect yourself with small samples. We deliberately name no target values here, because they depend on your archive and your risk.
| Metric | How You Collect It | Warning Sign |
|---|---|---|
| Share of statements with a citation | Go through a handful of answers per week and count statements without a number | Statements without a number pile up |
| Hit rate of the citations | Sample: does the passage really support the statement? | Citations that fit only by topic |
| Share of answers with “not in the documents” | From the history or from samples | Close to zero is suspicious, very high points to gaps in the archive or a search that is too narrow |
| Errors found per sample, by kind of error | Review log with the four kinds | A cluster in “outdated” points to a maintenance problem in the archive |
Splitting by kind of error shows where to start. Invented statements call for a stricter rule on sources and gap messages. Outdated statements call for maintenance of the archive. Misattributed statements call for unambiguous identifiers in the documents and some practice at comparing. Statements from another context call for training: whoever reads paragraphs in context catches them.
What You Can Take Away
If you are evaluating or already using an AI tool for company knowledge, these questions help:
- Does every statement carry its own citation, or is there only a list at the end?
- Can you open the document with one click and see title, number and date?
- Does the tool say “not in the documents” when nothing suitable exists? Test it with a question whose answer is demonstrably in none of your documents.
- Can you set where answers may come from, and does the permission of the asking person affect the search?
- Do your employees know the four kinds of error and the review depth for their use?
- Are version and date of your documents maintained, and who takes care of gaps?
Our position: invest first in citations, honest gaps and a review routine, and only then in the model. A better model helps, but it makes none of these three measures unnecessary. An assistant that errs now and then and cites every statement is easier to run than one that rarely errs and never shows where it got something.
If you would like to see how answers with sources and the choice of sources look in daily work, we are happy to show you in a demo, without obligation and with examples from your industry.
See Thero live
Book a short demo. You talk directly to the founding team.