Picture a steering meeting where the company has to decide on an enterprise AI. The head of IT says: “Open models only, and on our own servers.” The head of a business unit pushes back: “I want the best model available, no matter where it runs.” The data protection officer says one sentence: “As long as nothing leaves Europe.” All three have a legitimate concern. Still, the meeting ends without a decision, because each of them is answering a different question.
This article sorts those questions. Our thesis: Model choice and hosting are not matters of faith but several separate decisions that can be made with a handful of criteria: location, quality per task, speed, cost and replaceability. A second, practical thesis follows from it: for the tasks of a company there is rarely one model that is the right choice everywhere. If you choose per task and build the platform so that a model can be replaced later, you avoid two mistakes at once: paralysis caused by matters of principle, and dependence on a single decision.
A note up front. This article does not rate vendors or individual models, and it does not quote benchmarks. The figures that appear are round assumptions in clearly labelled example calculations. They are neither measurements nor prices. Where law comes into play we stay general: this is not legal advice, so please settle the individual case with your legal counsel.
Why the Holy War Asks the Wrong Question
The argument “open versus proprietary” mixes four questions that do not have to depend on each other:
- The task: what should the model do, and what happens if it gets it wrong?
- The model: who built it, are its weights (the numerical values of the trained model) available, and how large is it?
- The operation: who runs the model on which hardware, you, your platform vendor or a third party?
- The location: where does the data sit, and where is it processed?
These four questions are largely independent. An open model can run with an operator in Frankfurt or in a data center in a third country. A proprietary model may or may not be reachable through an interface inside the EU. Anyone who says “open” and means “secure” has confused two questions. So has anyone who says “proprietary” and means “our data is gone”.
The order is not arbitrary. Location, with everything that law and contract require of it, is a must-have criterion: it rules out options before you compare them. Only within what is permitted do you choose by task, model and operation. If you reverse the order and look for the “best” model first, you will later face the question of whether you may use it at all.
| Decision | Guiding question | Typical options |
|---|---|---|
| Location (must-have) | Where does the data sit, where is it processed, who has access? | Data storage in Germany, AI processing in the EU, other regions |
| Task | What should get done, and what does a mistake cost? | Triage, answers with sources, drafts, summaries |
| Model | How large, under which license, how is it updated? | Open or proprietary, small or large |
| Operation | Who runs it, who carries updates and availability? | You, platform vendor, partner |
A word on language: when people say “hosting” they usually mean two things at once, namely running the platform and running the model. Both can sit in different places and with different companies. That is exactly why it pays to look at them separately.
Open and Proprietary Models: What the Difference Really Means
Terms first. With an open model, the weights are available. You can download the model and run it yourself or with an operator of your choice. “Open” does not describe a single license, though. Some models may only be used commercially under conditions, and some restrict the purpose of use or the size of the company. Read the license of the specific model before you plan. With a proprietary model, you reach the model only through the interface of its maker or of an operator the maker has approved. The weights stay with the maker.
What does that mean in practice? Open weights give you freedom in operation and in version. You can decide where the model runs, and you can freeze a version so that answers do not change without your involvement. In return, you or your operator carry the hardware, updates, monitoring and availability. Proprietary models take operating effort off your hands and are often available earlier. In return, location, version changes and service life follow the maker’s offering.
Neither side is safer by itself. Security comes from contract, location, access control and operation, not from the license type. An open model on a poorly secured machine is less secure than a proprietary model with contractually agreed processing in the EU, and the reverse can just as well be true: a well-run open model may be the better choice. The license type only tells you which forms of operation are possible at all.
| Form of operation | Control | Effort on your side | What to look at |
|---|---|---|---|
| Open model, run by you | high | high | Hardware, updates, monitoring, skilled staff, the model’s license |
| Open model, run by an operator | medium to high | medium | The operator’s location and subprocessors, contract, freezing a version |
| Proprietary model through an interface | lower | low | Region of processing, storage of inputs, version changes |
| Proprietary model under a special agreement | medium | medium to high | Only if the maker offers it, check cost and term |
The table is a guide, not a verdict. Which row suits you depends on how much operating effort you want and are able to carry, and on what your location criterion allows.
A word on where TheroAI stands: we are a managed offering. We run the platform, data storage is in Germany and AI processing happens in the EU. Which partners are involved and which commitments apply contractually is set out in our data processing agreement and on the security page. This article is meant to help you ask the same questions of every offer, including ours.
Data Residency: Where the Data Sits and Where the Computing Happens
“Data in Europe” sounds like an answer, but it is usually shorthand for two different things. Data storage describes where documents, search indexes, conversation histories and logs are kept. Processing describes where the model reads the text excerpts and computes the answer. Both matter, and both can be in different places. Our wording for this is: data storage in Germany, AI processing in the EU.
Why separate them? Storage is permanent and large: it holds everything your company files as knowledge. Processing is momentary and narrow: the model only receives the excerpts needed for one question. An assistant should pass on only those passages that the person asking would be allowed to see anyway. That check belongs before the model call and not in the good manners of the model. How this looks in practice is described in our article on permission mirroring.
The following table shows what a statement about location answers and what it does not.
| Statement | What it answers | What it does not answer |
|---|---|---|
| Data storage in Germany | Where documents, indexes and logs are kept | Where a model processes the excerpts |
| AI processing in the EU | In which legal area the models run | Who has access in operations or support, whether inputs are stored |
| Contractually agreed zero retention | That inputs are not to be stored or used for training | How you can verify it, which takes a contract and evidence |
| Data processing agreement under Art. 28 GDPR | Roles and duties of controller and processor | Whether your purpose is permissible in the individual case |
The table leads to a conclusion: a statement about location is a good start but not a review. It does not replace questions about access, storage and contract. These six questions are worth putting to every operator that processes your texts:
- In which countries and with which subprocessors do the models that process your texts run?
- Are inputs or answers stored, and if so, for how long and for what purpose?
- Is your data used for training, and is the answer in the contract rather than only on a web page?
- Who has access to content in operations and support, and is that access logged?
- How do you learn about a change of operator or subprocessor, and what time limits apply to your objection?
- What happens to your data and indexes when the contract ends?
A legal frame belongs here too, though only in outline. If a vendor processes personal data on your behalf, a data processing agreement under Art. 28 GDPR is generally required. If an operator or its subprocessors are based outside the EU or have access from there, additional questions about transfers to third countries arise, and you should have them reviewed legally. This is not legal advice, so please settle the individual case with your legal counsel and your data protection officer.
Treating location as a hard must-have criterion saves time: from the list of conceivable models and operators, everything that does not meet the criterion drops out at the start. That shortens the comparison considerably, because you only choose within what is permitted.
Quality per Task: Measure Instead of Believe
The most important question of model choice is not “Which model is the best?” but “Which model is good enough for this task?”. Tasks in a company differ considerably in what they demand. Sorting a request into the right queue is a different job from checking a contract for deviations. The graphic shows an illustrative assessment of such requirement profiles. It is an assessment of what the task demands, not a measurement of models.
A fictional example from our demo documents makes this tangible. A customer writes: “The filler FL-200 reports error E-217.” That single request becomes three tasks:
- 1.Triage: does the request belong to service, and how urgent is it? The output is short, and a mistake is noticed quickly because a person sees the queue anyway. Here speed counts.
- 2.Look up and answer with sources: what do the maintenance report and the manual say about E-217? Here faithfulness to the source counts, because an invented remedy would be worse than no answer. An assistant should back every statement with a source and say “not found” when nothing is on record.
- 3.Draft the reply: tone, completeness and reference to the source. Before sending, the draft goes to a person for approval, so time pressure is low.
The same request therefore calls for three different profiles. Whether one model or several are chosen for it is not a worldview but a question of requirement and effort. The page on AI for plant engineering shows how such questions are answered with similar examples.
How do you know whether a model is good enough for a task, though? Not from public rankings. They measure general tasks on other people’s data and say little about whether a model copes with your maintenance reports, your technical terms and your document structure. A test set from your own daily work is more informative. Here is how to proceed:
- 1.Collect real tasks. For each use case, start with 30 to 50 cases from daily work, together with the documents they need.
- 2.Define in advance what a good answer is. Correct, backed by a source, in the desired format and, where nothing is on record, explicitly “not found”.
- 3.Have experts rate the results blind. Without knowing which model answered.
- 4.Count the types of error, not just the total. Invented details, missing sources and wrong triage weigh differently.
- 5.Repeat the test at every model change. Larger version jumps of the same model count too.
The figure of 30 to 50 cases is a practical starting point, not a statistical guarantee. It is enough to spot rough differences and small enough that the check actually happens. More important than the size of the test set is that it comes from your daily work and that you keep it, because it becomes the basis of every later model change.
Latency: What the Waiting Time Is Made Of
“Fast” is the quality that employees feel most. Two different times are meant by it: the time to the first word, which decides whether a chat feels responsive, and the time to the complete answer, which matters mainly for long texts and for processes running in the background. The time is made up of three parts: the path and the search, the start of the model, and the generation of the answer.
The following example calculation shows how these parts behave. All values are round assumptions for illustration, not measurements and not properties of particular models. The unit in which models process text is called a token, roughly a fragment of a word.
| Assumption (example calculation) | Small model | Large model |
|---|---|---|
| Path through the network inside the EU | 0.05 s | 0.05 s |
| Search and permission check (only for answers with sources) | 0.7 s | 0.7 s |
| Model start until the first word | 0.3 s | 1.0 s |
| Generation speed | 100 tokens per second | 40 tokens per second |
| Length of a triage result | 20 tokens | 20 tokens |
| Length of an answer with sources | 400 tokens | 400 tokens |
This gives four cases:
| Task | Model | First word after | Complete after |
|---|---|---|---|
| Sort a request | small | 0.35 s | 0.55 s |
| Sort a request | large | 1.05 s | 1.55 s |
| Answer with sources | small | 1.05 s | 5.05 s |
| Answer with sources | large | 1.75 s | 11.75 s |
To check the arithmetic: for the answer with the large model, 0.05 plus 0.7 plus 1.0 seconds pass until the first word, which is 1.75 seconds. Then 400 tokens are generated at 40 tokens per second, which takes 10 seconds, so 11.75 seconds in total.
Three things stand out. First, for the long answer, generation makes up most of the time, not the trip to the data center. Even if the path through the network were ten times as long in this calculation, 0.5 instead of 0.05 seconds, the total for the large model would only rise from 11.75 to 12.2 seconds. Inside the EU, speed is therefore decided mostly by model size and answer length, hardly by location. Second, when answers appear step by step, the first word matters most for how fast it feels. At 1.05 or 1.75 seconds that is a clear difference but not a break. Third, for short tasks such as triage, the start of the model dominates. A small model is not only cheaper here but also noticeably faster.
What follows for the choice? Distinguish interactive tasks from downstream ones. A chat in which someone is waiting for the answer calls for a quick first word. A background process that, say, prepares a draft reply and then waits for approval can afford a slower, more careful model, because the approval takes time anyway. Also do not calculate only the ideal case: under load, with queues and with long documents, the times change. Measure with your test set under realistic conditions.
Cost: Choosing per Task Instead of One Model for Everything
Not every request needs the same model, and models differ in price. That creates leverage: if a large share of requests is simple, that share can be handled by a smaller and cheaper model, and the large model is kept for the demanding cases. Again, an example calculation with round assumptions shows how big the leverage can be. The price factors are not prices but calculation values that illustrate the principle.
| Tier (example calculation) | Share | Requests per month | Price factor per request | Calculation units |
|---|---|---|---|---|
| Simple: small model | 70 percent | 7,000 | 1 | 7,000 |
| Medium: mid-sized model | 25 percent | 2,500 | 4 | 10,000 |
| Demanding: large model | 5 percent | 500 | 10 | 5,000 |
| Total | 100 percent | 10,000 | 22,000 |
For comparison: if each of the 10,000 requests runs on the large model, that is 10,000 times 10, or 100,000 calculation units. Choosing per task gets by with 22,000, which is 22 percent. If everything runs on the small model, it is 10,000 units, but then the demanding cases also get a model that was not meant for them.
How robust is this calculation? It rests on two assumptions, and you should check both before deciding. The first is the distribution of requests. Whether 70 percent of your cases are simple only your own stock shows, for example a sample of real requests. The second is the triage itself: it needs a method that recognizes which tier fits, and that method can be wrong. A stress test: if every tenth simple request is rerun afterwards on the large model, 700 times 10, or 7,000 units, are added. That makes 29,000 units, or 29 percent of the reference value. Choosing per task still stays clearly cheaper in this calculation, but the lead is smaller than in the first case.
Two caveats belong here. First, model costs are only one part of the total cost of an introduction. Integration, data preparation, training and operation come on top and are often larger than the model line. You can try a rough calculation of the benefit in your own company with the ROI calculator. Second, several models mean more administration. Every additional tier has to be tested, monitored and covered by contract. Start with two tiers, not five, and add more only when the test set shows a benefit.
Replaceability: The Model Will Change, the Process Should Stay
A thought experiment for your next steering meeting:
The model you choose today is discontinued in two years, becomes much more expensive, or is replaced by a better one. What would you have to change, and how long would it take?
If the answer is “everything”, you have committed yourself although you only wanted to choose a model. Dependencies arise in several places, and not all of them are obvious:
- Instructions: wording written for one model behaves differently with another.
- Search index: the number vectors used to find documents come from an embedding model. If that model changes, you generally have to index again, which costs time and computing effort.
- Output formats: tables, outlines and source references have to appear in the new environment in exactly the same way.
- Contracts: term, notice periods and the question of the format in which you get your data back.
- Habits: people have grown used to the answer style and the speed.
The most effective countermeasure is architectural: everything that has to be reliable belongs in fixed rules of the platform and not in the behavior of the model. Permissions, source references and approvals should hold no matter which model happens to be answering. In TheroAI they are fixed parts of the platform, and the model is the replaceable layer beneath. The approval step in the workflow editor is an example: title, approver and a deadline of 72 hours are set on the step, regardless of which model wrote the draft.
You will find more on such processes on the page about workflows. For the model change itself, a four-step approach is advisable that keeps quality per task in view: secure the test set, run the new model in parallel, compare per task and only then switch step by step.
The following table shows what stays the same in a change and what you should check again.
| Area | Stays the same | Check again |
|---|---|---|
| Permissions and sources | Fixed rules of the platform | Whether source references still appear in full with the new model |
| Approvals | Step, deadline, approvers | Whether drafts match the previous standard in tone and format |
| Search index | Unchanged as long as the embedding model stays | Indexing again if the embedding model changes |
| Logging | Audit log of the runs | What the new operator logs on its side |
| Instructions | Task descriptions | Test the wording again with the test set |
One point belongs in the discussion with the works council, if you have one: a model change can alter what is logged and who can evaluate what. Under section 87 (1) no. 6 of the German Works Constitution Act (Betriebsverfassungsgesetz), the works council has a right of co-determination on the introduction and use of technical devices intended to monitor the behavior or performance of employees. Whether that applies in the individual case is for your legal counsel to clarify.
Five Typical Mistakes in Choosing a Model
Five patterns can be derived from the sections above, and they come up again and again in practice:
- The largest model for everything. It is convenient and, for simple tasks, needlessly expensive and slow.
- The smallest model for everything. It saves the most and falls short on the demanding tasks, usually when that hurts most.
- Location checked only on the website. A statement on a page replaces neither a contract nor the question of access and storage.
- Rankings instead of your own tests. General comparisons say little about your documents and your technical terms.
- No way back planned. Whoever does not know how a change works has already committed.
The same pattern lies behind all five: a decision is treated as final although it does not have to be. The remedies are the same too, namely your own test set, a clear must-have criterion and a described path for changing models.
Our Position
Choosing a model is an engineering task with criteria, not a question of allegiance. We think it is right to fix the location first, then decide per task, and to treat the model as a replaceable layer. Then it is no longer dramatic whether an open or a proprietary, a small or a large model is answering today, because permissions, sources and approvals apply anyway. Whoever builds it this way can choose models by quality, speed and cost and change them later without reinventing processes.
We know that the model landscape changes quickly. That is exactly why we value fixed conditions more than any snapshot: a location criterion, a test set from your own daily work, and a path for change.
What You Can Take Away
Eight questions you can already ask in the next meeting:
- Have we set location as a must-have criterion, separated into data storage and processing?
- Do we know in which countries and with which subprocessors our texts are processed, and is it in the contract?
- Have we listed our tasks and noted for each what a mistake costs?
- Do we have a test set of real cases with which we can compare every model?
- Do we know what response time is made of, and do we distinguish interactive from downstream tasks?
- Have we calculated with our own distribution of requests instead of a flat value?
- Do permissions, source references and approvals apply independently of the model?
- Do we know how a model change works, including re-indexing, contract and consultation with the works council?
If you would like to go through the questions with a concrete use case, a demo is a calm opportunity to see TheroAI with your own examples and to discuss the open points. No rush: you can also answer the questions on your own first.
See Thero live
Book a short demo. You talk directly to the founding team.