Thursday, shortly before the shift starts. A maintenance technician at a bottling plant asks the assistant: “The filler FL-200 reports error E-217. Has this happened before, and what helped?” The answer arrives within seconds, cleanly worded, with two sources: the maintenance manual and an older fault report. It sounds complete. It is not.
This example is explicitly made up, it comes from the fictional demo documents on our industry pages, but it is typical. What the assistant is missing is the service report from the last visit. It sits as a photo PDF called “Scan_0048.pdf” in a folder, contains the service technician’s handwritten notes, a company stamp, and not a single letter a program could read. That is where it said that for E-217 a sensor was replaced and not the seal. The assistant did not lie. It answered with what it could read. That is the real problem.
Our thesis in one sentence: before you connect a knowledge system, check what in it is readable, and demand from every system that it reports what it cannot read instead of silently moving past it. This article shows where documents become unreadable for an AI, how to measure that with a small sample, which cleanup work is worth the effort, and what an assistant should make visible on its own. What this looks like in day to day production is described on our page AI for Manufacturing.
The Silent Gap: Why Gaps Are More Dangerous Than Errors
A system that fails loudly is easy to manage. It shows an error message, somebody takes care of it, and an hour later it works again. A knowledge assistant fails differently. It answers from what is there, and what is not there leaves no trace. Three properties make that treacherous:
- A document that cannot be read produces no error message in the answer. It is simply missing.
- The answer stays fluent and confident, because language models build a rounded text even from thin material.
- Whoever checks the sources under the answer checks the sources that are there. The missing source appears on no list.
We call this the completeness illusion. It does not come from ill will in the technology but from the nature of the thing: an answer can only report on what is searchable. About the unsearchable it can say nothing, not even that it exists.
Thought experiment: your archive holds 1,000 documents, and 140 of them are invisible to the AI. How would you notice with the next answer that it is 140 and not zero?
The honest answer is: you would not. That is exactly why the check of data quality belongs before the connection and not after it, when the first answers are already in circulation. The number 140 is not a measurement here but the worked example we build further down in this article.
The Path of a Document: Four Stations Can Lose Something
To see where a document becomes unreadable, it helps to look at the path it takes. In most knowledge systems it is similar, even if the terms differ: the document sits in an archive, gets read, is cut into sections, is searched for fitting passages when a question comes in, and the answer is built from those passages.
At each of the first four stations a different kind of information is lost:
- Filing: A name like “final_v2_new” says neither what the document contains nor whether it is valid. Several versions sit side by side without any marking.
- Reading: A scan without a text layer has nothing that could be extracted. Stamps and handwriting stay empty or are recognized wrongly.
- Splitting: A table row loses its column headers, multi-column text is read across, and sentences appear that nobody ever wrote.
- Finding: The old version is just as findable as the new one, because both contain the same wording.
The fifth station, the answer, is the only one you see. It turns the gaps of the other four into one closed, readable text. The gaps therefore arise before the AI sees anything, and that is the good news: a large part of them can be fixed by tidying up on the side of the documents, without swapping a model.
Six Kinds of Documents an AI Reads Badly or Not at All
In practice the problems boil down to six patterns. The table below summarizes them, and the sections after it explain each one.
| Kind | What happens | How you recognize it | Remedy |
|---|---|---|---|
| Scan without a text layer | The document is only an image, there is no text to read | You cannot select a sentence in the PDF | Text recognition, or a fresh export from the source |
| Table | Rows lose their column headers, merged cells slip | Numbers are quoted without unit or context | Clear header row, one table per sheet, spreadsheet instead of PDF |
| Multi-column layout | The text is read across the columns | Quotes break off mid sentence | Single column version, clean export |
| Stamp and handwriting | Margin notes and stamps are missing or misread | The decisive remark is the one that is missing | Capture what matters as text |
| File names and filing | The name reveals neither content nor status | Names like “final_v2_new” or “Scan_0048” | Naming rule: type, topic, ID, date |
| Outdated versions | Old and new versions sit side by side with equal standing | Two answers that contradict each other | One valid version, archive the old ones |
How reliably a kind is read is a matter of degree and not of yes or no. The overview below places typical document types. It is a rule of thumb from practice and not a measurement.
Scans Without a Text Layer: The Picture of a Text Is Not Text
A scanner first delivers an image. Only text recognition lays an invisible text layer over that image. If it is missing, a program sees nothing but pixels. To you the page looks readable, to a system it is empty. The simplest test takes ten seconds: try to select a sentence in the PDF and copy it. If that does not work, there is no text layer.
Even later text recognition is not flawless. It depends on the resolution of the scan, on skew, stains and typeface. A phone photo with the page lying crooked in the light is a harder case than a clean scan at 300 dots per inch. Models that understand images can often read more than classic text recognition. But they need more time and computing effort, and they do not read without mistakes. For important documents it is therefore better to rely on a clean source: if the document also exists digitally, export it again as a PDF from the original program instead of scanning the printout.
Tables: Numbers Without Context Are Worse Than No Numbers
A table is only knowledge if every cell knows what it belongs to. In a fictional acceptance report from our demo documents, the row “Format change” holds the value “maximum 30 minutes” under “Agreed” and “34 minutes” under “Measured”. If the row loses its header, a 34 is left over. Whether that means minutes, bottles or percent, nobody knows any more, not even the AI, and it guesses.
Typical stumbling blocks are merged cells, subtotals in the middle of a table, several tables on one spreadsheet sheet, hidden columns, and header rows that run over two lines. On top comes a classic in German speaking countries: the string 1.249 means one thousand two hundred forty nine to us, and one point two four nine in other countries. Where numbers in German format land in systems with a different convention, quiet errors arise. So always check one number against the original when it comes to tables. And whenever possible, keep tables as spreadsheet files instead of as PDF printouts of a table.
Multi-Column Layouts: Reading Across Builds Sentences That Do Not Exist
Datasheets, standards, brochures and product catalogs are often set in two or three columns. A human follows a column downward. A program that reads a page line by line jumps from the left to the right column and back, and a sentence appears that consists half of one column and half of the other. Headers and footers add to it: “Page 3 of 12, confidential” then sits between two paragraphs and is treated like running text.
The test is simple: copy two paragraphs from the PDF into a text editor and read them aloud. If what appears there sounds like the document, the page reads well. If the sense jumps, the layout is a problem. Modern reading methods recognize columns, but not every layout works, and the more unusual the structure, the more it comes down to a sample you have to take yourself.
Stamps and Handwriting: The Most Important Note Is in the Margin
The incoming stamp with a date, the stamp “approved”, the handwritten margin note “do not replace seal, sensor”: the very pieces of information that separate a document from a draft are often outside the printed text. For any reading method they are the hardest class. Handwriting is sometimes hard to decipher even for humans, stamps overlay the text beneath them, and signatures are not language anyway.
So do not count on such entries being read reliably, and do not let any decision hang on them alone. If a stamp or a note carries meaning, capture it additionally as text: in a field of the business system, in the file name, or in a short accompanying note. A remark on data protection: handwritten notes and signatures can contain personal data. Whether and how you may process such documents depends on the individual case. This is not legal advice, so clarify it with your legal counsel or your data protection officer.
File Names and Filing: The Name Is the First Piece of Information
File names are underrated because people have the context in their heads. Whoever has been with the company for a long time knows that “Scan_0048” is the service report for the filler. For everybody else, and for every system, the name is the first and sometimes only piece of information about a document before its content is even read. A name that states type, topic, ID and date helps with searching, with checking the source and with recognizing the valid version. The examples are made up, the principle is not:
| Before | After | What changes |
|---|---|---|
| Scan_0048.pdf | Servicebericht_FL-200_E-217_2025-03.pdf | Type, machine, error and month are in the name |
| final_v2_neu.docx | Wartungsplan_FL-200_gueltig-ab-2025-01.docx | The status is recognizable |
| Tabelle1.xlsx | Ersatzteile_FL-200_Preise_2025.xlsx | Content and year are clear |
| Vertrag_alt_NICHT_NUTZEN.pdf | Archiv/Rahmenvertrag_TP-118_2022.pdf | The folder says “superseded”, not the file name |
The file names in the table stay in German on purpose, because they are the example documents of our fictional demo. A naming rule is the cheapest measure in this article, and it works without any AI: your colleagues find what they are looking for faster.
Outdated Versions: The Old Version Is as Findable as the New One
The problem of outdated versions is invisible to humans because they know the context. For a search system, two documents are simply there, both fit the question, both contain similar sentences. Without a date, a version number or a note about a replacement there is no reason to prefer one over the other. The answer then mixes versions or takes the wrong one. Prices, deadlines and procedure instructions are especially delicate, since over the years little changes in the wording but a lot in the number.
The remedy is twofold: put superseded versions into an archive area that you do not connect, and have every answer show you which document and which version it rests on.
Measuring Data Quality Before You Connect Anything: The Sample
Most companies know their archive by feel and not in numbers. That is normal, because nobody ever needed what an AI needs now: an overview of how many documents are machine readable. The good news is that you do not need a full audit. A small random sample gives you the order of magnitude. The following worked example is explicitly illustrative. It shows the method, not your numbers.
The Method in Three Steps
- 1.Draw 50 documents at random from the archive you want to connect. Random means: not the prettiest, not the most important, but roughly every twentieth one.
- 2.Check each document with the four questions of the cleanup checklist further below and note exactly one finding: no finding, scan without a text layer, table with a problem, columns mixed up, or outdated version.
- 3.Extrapolate: if the archive has 1,000 documents and you check 50, multiply by 20.
| Assumption in the worked example | Value |
|---|---|
| Size of the archive | 1,000 documents (assumption) |
| Sample | 50 documents drawn at random, which is 5 percent |
| Extrapolation factor | 20 |
Suppose the sample gives the following picture:
| Finding | In the sample | Extrapolated to 1,000 | Share |
|---|---|---|---|
| No finding | 34 | 680 | 68 % |
| Scan without a text layer | 7 | 140 | 14 % |
| Table with a problem | 4 | 80 | 8 % |
| Columns mixed up | 3 | 60 | 6 % |
| Outdated version | 2 | 40 | 4 % |
| Total | 50 | 1,000 | 100 % |
The example carries two messages. First: 320 of 1,000 documents, so 32 percent, need attention, almost a third. Second: the largest single item is the 140 scans without a text layer, and those are exactly the ones users cannot see. A sample of 50 documents tells you the order of magnitude, not the second decimal place. If in practice you find 14 percent scans, or 10, or 20 percent, that rarely changes the order of the work. It shows whether you have a side issue or a main issue.
The Second Check: Questions With a Known Answer
The sample checks documents. The second check tests the result. Collect ten to twenty questions whose correct answer you know and which really occur in your day to day work, for example: “What notice period applies in the framework contract with TP-118?” or “What was the cause of E-217?”. Put them to the assistant and compare. Every wrong or incomplete answer has a cause in the documents, and usually it is one of the six kinds from this article. These test questions are also your regression check: after every cleanup and after every new connection, you ask the same questions again.
The check with questions is more honest than any statistic, because it measures exactly what counts: whether the people who ask get correct answers.
The Small Cleanup Checklist: Four Questions per Document
Cleaning up is not a big project as long as you limit yourself to what your people actually ask about. Start with the twenty topics that come up most often in daily work, and run the documents belonging to them through these four questions:
- 1.Can you select a sentence? If not: run text recognition or export the document again from the source.
- 2.Is it the valid version? If not: archive the old version and keep the status in the name.
- 3.Does the file name say what is inside? If not: rename it following the rule type, topic, ID, date.
- 4.Is the table still a table? If not: keep it as a spreadsheet file or add a header row.
For the archive as a whole, five habits help that cost nothing but prevent a lot:
- Every folder has a responsible person who says what is valid there.
- New scans land in an intake folder and only move into the archive after text recognition.
- Superseded versions move to an archive area that is not connected.
- There is a single naming rule, short enough that everybody remembers it.
- The sample is repeated after the first run to see whether the findings have changed.
This list is short on purpose. Whoever tries to clean up everything first will never begin. Whoever starts with the most frequent questions has, after a few weeks, the documents in order that are needed every day.
What an Assistant Should Report Instead of Failing Silently
So far this was about your documents. The second half of the thesis concerns the system: a knowledge system should make the unreadable visible. If a document cannot be read, or only partly, it needs a state that people can see, not a silent gap. That is not a nice extra but the precondition for cleanup to work at all: whoever does not see which documents drop out cannot fix them.
In TheroAI every connected document has an indexing status. It appears in the document’s details, and the per app statistics count failed, not started and unsupported documents separately. The most important messages of the interface, shown here with their English labels, and what you do about them:
A few notes on these:
- Partially indexed means, according to its hint text, that the document exceeds the upper limit and only the first part is searchable. That is an honest statement about a loss you would otherwise never have noticed, and the solution is usually to split a long document into sensible parts.
- No content means, as we understand it, that reading produced no text. A scan without a text layer is a typical candidate here, and depending on the setup it ends up at “Needs multimodal model” instead. An empty file belongs here as well.
- Not supported reports that the file type is not processed. That is inconvenient and still better than silence, because you know there is a hole.
- Needs multimodal model appears, as we understand it, when a document can only be read with a model that understands images, and none is set up. For scans and photos this is the switch between empty and readable.
- Besides that, the product points out when individual pages were read from the text layer only, and tables or figures on those pages exist as plain text.
At the level of the connection you see the balance at a glance. In the demo environment of our capture, the list of apps shows for Google Drive 396 documents, 297 of them indexed, and for Gmail 6,743 documents, 5,915 of them indexed.
The difference, 99 documents for Drive and 828 for Gmail, is not yet a finding. Some documents wait in the queue, others are being processed right now. But it is the list you work through: look at which statuses hide behind it, and you know whether it is a lead time issue or a quality issue. How the connection of apps and connectors works in detail, we have described elsewhere.
What a Status Does Not Do
Honesty belongs in the description. A status says whether a document was read, not how well. A document with the status “Indexed” can still contain mixed up columns, a table without context or an old version. The status closes the coarsest gap, the silent one. It replaces neither the sample nor the questions with a known answer. And it does not change the fact that the assistant shows numbered sources under every answer that you should open in spot checks: whoever never reads the source relies on the very system they wanted to check.
What Cleaning Up Cannot Do
Clean does not mean correct. A perfectly readable document can be wrong in content, and a neatly named contract can have been outdated for a year without it showing in the name. Cleaning up makes sure the AI sees what you want to show it. Whether it is true is still decided by your expertise, which is why a responsible person per folder is one of the five habits.
The work also has limits in effort. Handwritten legacy files, microfilm and printouts from the nineties cannot be fully prepared, and you do not have to. Decide deliberately what gets connected, and record what does not: a short list “deliberately not connected” is itself a piece of data quality, because it prevents anyone from assuming the AI knows about it.
And finally there is data protection: documents with personal data, such as personnel files, applications or scanned signature lists, are not suitable for connection merely because they are readable. Whether and under which conditions they may be processed is a question for the individual case. This is not legal advice, so clarify it with your legal counsel and your data protection officer before you connect.
What You Can Take Away
An AI is only as well informed as the documents it can read. Whoever connects without checking gets a system that answers with full confidence on a part of the knowledge, and in our worked example that would be the 68 percent of the archive with no finding. The rest is the silent gap that gives this article its name. Our position: check first, clean up what is needed often, and demand from every system that it reports what it cannot read.
Six check questions for next week:
- Have we checked 50 random documents from the archive and counted the findings?
- How many scans without a text layer are among them, and do they sit in topics people ask about often?
- Have we collected ten to twenty questions with a known answer that we use to test every change?
- Is there a naming rule and an archive area for superseded versions?
- Does the system visibly report unreadable, partly read and unsupported documents, and who looks at these messages regularly?
- Who is responsible per folder for making sure that what is valid is what is stored there?
And three steps for the week after: draw the sample and extrapolate. Clean up the documents for the twenty most frequent topics. Only then connect, and repeat the sample after the first run.
If you would like to see how such a connection looks with status messages and sources in a demo environment, we are happy to look at it together. A short demo is enough, and afterwards you can decide in peace what the next sensible step for your archive is.
See Thero live
Book a short demo. You talk directly to the founding team.