Picture a Friday afternoon. The demo went well: twenty sample documents, a few clean questions, precise answers with sources. On Monday the whole thing is supposed to be “just” plugged into the real file share, the engineering SharePoint and the sales team's mailboxes. Three weeks later only part of the data is searchable, the data protection officer has follow-up questions, and nobody can say whether the answers are complete.
This scenario is a thought experiment, not a customer case. In this form or a similar one, though, it shows up in almost every project where AI is meant to meet company data. Our thesis: connecting the systems is the real project. The language model can be swapped and shown off within days. The connection, by contrast, is ongoing operations. It has a sign-in, a first large run, continuous updating, deletions, outages, and source data that nobody prepared for an AI.
This article walks through these six stations in order. It calculates with an example how long a first sync can take, and uses the connector catalog of TheroAI with its 29 apps to show what we mean by a visible state. At the end you will find a position and a list of questions you can put to any vendor, us included.
Why the Demo Says Nothing About Operations
A demo answers exactly one question: can the system give sensible answers from these documents? It does not answer what comes afterwards. The difference fits into one table.
| Aspect | In the Demo | In Operations |
|---|---|---|
| Data volume | A few selected documents | Years of files across several systems |
| Access | One test account with broad rights | Many people with different rights |
| Freshness | Yesterday's state is fine | Changes, deletions and new versions pile up daily |
| Errors | Cleared out before the meeting | Appear at random and someone has to notice |
| Ownership | Whoever built the demo | IT, business units, data protection and vendor together |
The core is in the fourth row. In the demo there are no errors because somebody found them beforehand. In operations there always are, and the decisive question is not whether they occur but who notices and how quickly. A connector that fails silently is worse than one that fails loudly: the answers still look complete even though part of the corpus is missing.
Six Stations Where Time Gets Lost
Every integration goes through the same stations, whether the source is Google Drive, SharePoint, a ticket system or a wiki. The order is also the order in which the problems show up.
The stations are not equally expensive. Projects rarely lose most of their time to the technology itself. They lose it to coordination (station 1), to underestimated volumes (station 2) and to surprises nobody sees (stations 4 to 6). Let us take them one by one.
Station 1: Sign-In and Permissions
The first loss of time happens before a single document has been read. A connection needs access, and the person leading the project rarely grants it. For Google Workspace sources, the setup in TheroAI asks for the credentials of a service account and the address of an administrator. For Microsoft sources it runs through an app registration in Azure AD that an administrator has to consent to. These are not quirks of one vendor. They follow from the fact that sources rightly protect their data.
What matters is the distinction between two kinds of rights, which tend to get mixed up in conversations:
- The right of the connection: What may the connector read overall? This is an organization-level decision. It can be narrowed, for example to certain folders, file types, labels or channels.
- The right of the individual: What may a particular employee see in answers? This is a per-document decision, and it should be taken over from the source, not reinvented.
Both belong to the setup. In the setup of the Google Drive connector, TheroAI lists the three points that count as ticked lines: user permissions for documents are indexed along with them, content and permission updates are synchronized several times per hour, and real-time synchronization is supported. Below that comes the entry of the authentication data, with the setup guide on the right.
How many people need to have a say at this station is almost always underestimated:
- The source's IT administration, which actually grants the access.
- The business units that own the content and decide what may enter the search space.
- Data protection, which reviews the data processing agreement under Art. 28 GDPR.
- The works council, as soon as mailboxes, calendars or chats are to be connected. For technical systems that can be used to monitor behavior or performance, a co-determination right under section 87 (1) no. 6 of the German Works Constitution Act may apply.
- Management, which decides which sources are connected first.
This is not legal advice; clarify the individual case with your legal counsel. In practice: whoever convenes this group only after the technical setup loses weeks. Whoever plans it beforehand loses one meeting. You will find more on TheroAI's approach to access and data residency on the Security page.
Station 2: The First Sync
The first run reads the entire released corpus, breaks it up and makes it searchable. How long that takes depends on three things: the volume, the rate limits many sources place on requests, and the effort of reading and preparing each document. A twenty-page scan is not a text document with two paragraphs.
An example calculation makes the order of magnitude tangible. The numbers are assumptions, not measurements.
| Assumption | Corpus | Throughput per Hour | Duration | In Days |
|---|---|---|---|---|
| Slow | 200,000 documents | 2,000 documents | 100 hrs | just over 4 days |
| Medium | 200,000 documents | 5,000 documents | 40 hrs | about 1.7 days |
| Fast | 200,000 documents | 10,000 documents | 20 hrs | about 0.8 days |
The calculation is deliberately simple, and so is its message: depending on throughput, the same source can keep a system busy for half a day or for almost a week. And that is pure runtime, assuming no interruptions. If the source throttles, a restart happens or maintenance gets in the way, it takes longer.
More important than the number is the question of what users see in the meantime. While the run is in progress, only part of the corpus is searchable, and every answer comes from part of the truth. A system that does not show this creates trust it has not earned. That is why the app list in TheroAI shows, for each connection, the content indexing state (“Synchronisiert” or “Vorgang läuft” in the German interface) and the number of documents, together with the number already indexed.
A practical piece of advice follows from the calculation: do not start with everything. A partial corpus that you narrow down with filters is searchable within hours and shows you early which problems await you at stations 3 to 6.
Station 3: Ongoing Updates
After the first run the real work begins. Documents are edited, new ones arrive, permissions change. A sensible connector then does not read everything again but only what has changed since the last run. This is called delta synchronization. The question is when such a run starts.
| Strategy | How It Works | Strength | Price |
|---|---|---|---|
| Notification by the source | The source reports changes itself and the connector reacts | Very short delay | Only where the source offers it |
| Scheduled | Runs at a fixed interval and in a defined time zone | Predictable, even load | Delay until the next run |
| Manual | The run starts on a button press | Full control | Gets forgotten in daily work |
For Google Drive, the setup in TheroAI names two properties side by side: updates of content and permissions several times per hour, and support for real-time synchronization. Together this means that the normal case is a run at short intervals, and where the source reports changes, things go faster. We do not promise a fixed figure such as “after exactly two minutes” because it depends on the source.
There is no universal answer on which strategy to pick, but there is a usable rule of thumb: where permissions are sensitive, such as in HR or contract folders, a short interval or a notification from the source is worth it. Where content rarely changes, a longer cycle is enough and the load stays even. The manual run is no substitute but a sensible addition: after a larger reorganization of the filing, for example when whole folders were moved or shares were changed in bulk, it shortens the time in which the index still shows the old state.
Two things are easily overlooked at this station. First, the permission has to stay current alongside the content. A document that sits correctly in the index but is served to the wrong people is worse than one that is missing. Second, both directions matter equally: a newly granted access reaches the index only once the source reports it or the next run has taken place, and the same holds for a revocation.
Station 4: Deletions
Adding is easy, removing is hard. The index is a copy, and a copy knows nothing about the original having disappeared as long as nobody tells it. Three cases need to be told apart: a document is deleted in the source. A share is revoked, so the document stays but no longer for everyone. Or the entire connection is switched off.
For the first two, the same honest limit applies as for updates: the index is only as current as the last run or the last notification from the source. In between lies a window in which an answer can still point to a document that no longer exists in the source.
The graphic shows an example with an assumed interval of 20 minutes between two runs. The actual interval depends on the source and on the configuration. What matters is knowing that this window exists, and knowing it before audit or data protection ask about it. It is not a defect of one particular product but a property of every solution that copies content in order to answer quickly.
The third case, switching off, is explicitly designed in TheroAI: when a connection is deleted, all synchronized records and files, all indexed data including embeddings, and the configuration are removed. The dialog states that the action cannot be undone. That matters for deletion concepts, because deleting also includes the way back: whoever cuts off a source must know that no copy remains. How a deletion concept is to be shaped legally is something to clarify with your legal counsel.
Station 5: Outages
Sources are not always reachable. Credentials expire or are revoked because someone tidied up the account. A source throttles requests or is temporarily down. Individual documents cannot be read because they are damaged or too large. None of this is unusual, all of it happens.
| What Happens | Effect on Answers | How You Should Notice |
|---|---|---|
| Credentials expire or are revoked | Nothing new arrives anymore | Connection status, time of the last sync |
| Source throttles or is temporarily unreachable | Runs are delayed | A running job, missing progress |
| Individual documents are unreadable | Gaps in the corpus | Failed entries by type and reason |
The most dangerous outage is the silent one. If a connector has not read anything for two weeks and nobody sees it, everyone keeps answering with the state of two weeks ago. That is why a connection needs more than a “connected” tick. The overview of a connector in TheroAI shows the time of the last synchronization and the throughput of the last seven days. The “Protokolle” (logs) tab lists, under “Fehlgeschlagene Einträge” (failed entries), what could not be processed, grouped by type and reason. And two buttons allow the next step: “Neu synchronisieren” (sync again) and “Fehlgeschlagene neu indexieren” (reindex failed).
Station 6: Source Data Quality
The last station is the most uncomfortable because it lies not with the vendor but in your filing. An AI search makes visible what used to be bad in the dark.
A fictional example from the demo material of the Manufacturing page: someone asks how to resolve error E-217 on the filler FL-200. The filing holds two versions of the service manual, an old one in the wiki and a new one on the file share. Both are correctly connected, both are searchable. Which version the answer names is a question of filing, not of AI. And a scan of the machine file without a text layer is missing from every answer, without anything being reported as an error.
This does not call for a clean-up project before the integration that nobody ever finishes. A targeted approach works better:
- 1.Pick an area where questions are frequent and the filing is reasonably well kept.
- 2.Connect it and use ten real questions to see which problems the answers reveal.
- 3.Fix the causes in the source, not in the AI system.
- 4.Widen the area only once the answers hold up.
Transparency: What a Visible State Means
The Catalog
Six stations where something can go wrong lead to a simple demand on every vendor: the state of the connection must be visible without anyone having to ask. That starts with the catalog. TheroAI currently lists 29 apps that can be connected in read-only mode, grouped by provider such as Google Workspace, Microsoft 365 and Atlassian, and by category: productivity, communication, storage, development and documentation. You will find more on the scope on the Connectors page.
The Overview per Connector
More important than the number of apps is what you can see after connecting. The overview per connector lists documents by their state. The following graphic shows an example from the overview of a Google Drive connector in a demo instance.
The numbers belong to a demo instance, not to a customer: 396 documents in total, of which 297 are indexed (75 percent), 38 failed and 58 are not supported. The remaining three are in another state. So part of the documents is not searchable, and the overview says so openly instead of hiding it behind a green number. The states have clear meanings:
- Indexed: The document has been read and is searchable.
- Failed: Processing did not succeed. “Fehlgeschlagene neu indexieren” starts a new attempt, for all 38 in the example.
- Not supported: The file type is not read. This is not a malfunction but a limit you should know about.
- In progress and Not started: The document is on its way or waiting.
Such a display answers three questions that otherwise only get asked in a dispute: how much of the corpus is really searchable? What failed, and why? Which file types fall out? It does not prove that the answers are right, though. It shows what is searchable, not whether the content is correct. And it shows the state of the documents the connector knows about. What you deliberately exclude through filters does not appear there as a gap. The filter is a decision of its own that deserves documentation.
Our Position: Integration as an Operational Service
From all of this we draw four conclusions, against which we also measure our own work:
- 1.Integration is an operational service, not a setup step. It is not done once but run continuously, with someone responsible on your side and someone on ours.
- 2.Freshness is a number, not a feeling. Ask for the interval between runs, not for the word “synchronized”.
- 3.Errors have to be loud. A connector that fails silently is a risk to every answer.
- 4.Limits belong in the interface. Unsupported file types, failed documents and the window between two runs are not embarrassments but operational information.
This is less comfortable than the promise of a one-click integration. But it is the difference between a demo that looks good and a system you still believe in three months from now.
What You Can Take Away
Walk through the six stations with your vendor or your own project. These questions help:
- Sign-in: Who grants the access, and who has to be asked beforehand (IT, business unit, data protection, works council)? Can the scope be narrowed?
- First sync: How many documents does the corpus hold, and how long does the first run take under cautious assumptions? What do users see while it runs?
- Updates: How long is the interval between two runs, and where does the source report changes itself? Does that also apply to permissions?
- Deletions: How long can an answer still point to a deleted document, and what happens when a whole connection is removed?
- Outages: How do you see that a connection has not read anything for days? Where are failed entries listed with their reason?
- Quality: Which ten real questions test the first area before you widen it?
Add three steps for the next two weeks:
- 1.Pick an area with frequent questions and well-kept filing, and decide who is responsible for it.
- 2.Clear access, data protection and, where applicable, the works council before the technical work begins.
- 3.Decide which figures will tell you after four weeks whether the integration holds up: share of indexed documents, number of failed entries, time of the last synchronization.
If you would like to see how an integration with its status display feels in practice, we are happy to show you in a demo, calmly and with an example of your own.
See Thero live
Book a short demo. You talk directly to the founding team.