It is Monday, just after eight. The managing director of a machine builder with 140 employees bought an AI tool on Friday and sends a company-wide email: “An AI assistant is now available to us. Please use it.” Eight weeks later four people use the tool regularly, two have tried it once, the works council heard about the rollout over lunch, and when someone asks whether it is worth it, nobody answers with more than a feeling.
The example is made up. The situation will still look familiar to anyone who has rolled out new software without a frame around it. In stories like this the tool is rarely the problem. Something else is missing: a goal that can be checked, people with clear tasks, a schedule with decision points, and an agreement on what happens if it does not work.
This article describes such a frame for a mid-sized team: 13 weeks, so about 90 days, in three phases. It contains a week-by-week plan, the roles, an example calculation of the effort, the typical pitfalls and the criteria at which stopping is the right decision. Our thesis comes first so you can test it while reading: The goal of the first 90 days is a good decision, not a particular result. Anyone who promises results before a single pilot run has happened is selling. Anyone who plans decisions is working.
What 90 Days Can Deliver and What They Cannot
Thirteen weeks is short. It is not enough to change a company, and it is not enough to predict reliably how much time will be saved in the end. It is enough to learn on one use case whether the path is worth following, and to back that up with figures and feedback instead of gut feeling.
| Realistic after 90 days | Not honest to promise |
|---|---|
| One use case running in daily work for one group | A saving in hours or percent before it has been measured |
| Before and after figures for exactly that case | That every department benefits to the same degree |
| Rules for data, access and approvals, written down | That employees will adopt the assistant without guidance |
| A decision with reasons: continue, adjust or stop | That the investment pays for itself within a set time |
| A named person who owns operations | That the AI’s answers are reliable without review |
The right-hand column is not caution for its own sake. It describes what an honest plan cannot know at the time of planning. A roadmap that guarantees results is a brochure. A roadmap that guarantees decisions is a working tool.
The Roadmap at a Glance
The plan has three phases and three gates. A gate is a fixed date at the end of a phase on which a small group decides whether to continue, to improve or to end the effort. The gates fall on days 21, 63 and 91, which is the end of weeks 3, 9 and 13.
The split follows a simple line of thought. Three weeks of foundations are short enough that preparation does not become an end in itself, and long enough to settle data protection, permissions and the works council properly. Six weeks of pilot give the group time for the bumpy first days and still leave room to watch usage and quality over several weeks. Anyone who tests for only two weeks mostly measures curiosity. Four weeks of expansion are enough to onboard a second group and organize operations, and short enough that the final decision is not postponed.
Roles: Who Does What Over the 13 Weeks
An AI rollout rarely fails because of technology. It often fails because of responsibilities nobody spelled out. Five roles are enough for a mid-sized team. One person can carry several roles, but every role needs a name.
| Role | Task | Decides on |
|---|---|---|
| Sponsor (managing director or head of division) | Releases time and budget, attends all three gates, holds the line on decisions | Start, continue, stop |
| Project owner | Steers the plan, measures, draws samples, writes the report | Order and scope within the pilot |
| Pilot group (six people from the business unit) | Works on real tasks and gives honest feedback | Assessment in daily work |
| IT and data protection | Settles contracts, permissions and integrations | Which data and systems are connected |
| Works council (if there is one) | Is informed and involved from week 1 | Co-determination under the Works Constitution Act |
Two remarks. First, the sponsor is not an honorary role. Without someone who releases time when in doubt and accepts an uncomfortable result, the pilot gets done on the side and is judged accordingly. Second, the pilot group should be mixed. If only the enthusiasts take part, you learn that enthusiasts like the tool. You knew that before.
The Effort as an Example Calculation
How much time this costs cannot be said in general. The following is an example calculation with round assumptions that you should replace with your own. Above all it shows where the effort lies: not in the technology, but in the people who learn and check.
| Role | Assumption | Hours |
|---|---|---|
| Project owner | 6 hours per week, 13 weeks | 78 |
| Pilot group | 6 people, 2 hours per week, 6 weeks | 72 |
| IT and data protection | 12 hours in phase 1, 4 in phase 2, 8 in phase 3 | 24 |
| Sponsor | 1 hour per week, 13 weeks | 13 |
| Works council | Assumption: 6 hours in total | 6 |
| Total | 193 |
At eight hours per day, 193 hours are about 24 person-days. Not included are the onboarding of colleagues who join later and the running cost of licences. The calculation is only useful if you compare it at the end with the benefit you actually measured. In the ROI calculator you can play through your own assumptions without turning them into a promise.
Phase 1: Foundations (Weeks 1 to 3)
| Week | Focus | Result at the end of the week |
|---|---|---|
| 1 | Mandate, sponsor, roles | One sentence stating the goal, named people, released time budget |
| 2 | Choose the use case, measure the baseline | One use case with a review rule, times for typical tasks |
| 3 | Rules for data, access and approvals | Written rules, works council involved, gate 1 |
Week 1: One Sentence You Can Check
Start with a goal sentence that does not read “introduce AI” but describes something checkable. An example: “By day 63 we will find out whether our service team can handle fault enquiries faster and with verifiable answers using an assistant.” That sentence holds a team, a task, a date and a yardstick. It promises nothing, it schedules a test. The sponsor signs it, the pilot group is named, and the time budget from the example calculation is released.
Week 2: Choose the Use Case
Most pilot projects are decided on the wrong task. Either it is too rare, so no experience builds up, or it is too delicate, so every answer slows down the review. Four questions help with the choice:
- How often does the task occur? You need enough repetitions to see something after six weeks.
- What happens if there is a mistake? The further a wrong answer travels outward, the stricter the review has to be.
- Can the answer be checked against a source? Tasks whose answer sits in documents are easier to assess than tasks that depend on judgement.
- Are the documents available digitally, and may the pilot group see them? Without access there are no answers.
The first two questions can be combined in a four-field matrix:
The examples in the graphic are orientation, not a recommendation for your company. Place your own tasks. In manufacturing a typical candidate is looking up faults, for instance whether a particular error has occurred before and what helped back then. The page AI for manufacturing shows what that can look like. In the demo material it is an explicitly fictional example: the filler FL-200 reports error E-217. The task is frequent, a mistake can be spotted at the source, and the final decision on the repair stays with the service technician.
Just as important as the choice are the baseline values. In weeks 2 and 3, have the pilot group note for ten typical tasks how long they take and where questions arise. That costs little and is the only basis on which you can say in week 9 whether anything got faster. Without a baseline every statement stays “felt”.
Week 3: Rules Before Anyone Uses the Assistant
The rules need to be neither long nor legalistic. They must exist before the pilot begins. Six points belong in them:
- Which data may go into the assistant and which may not, for instance personnel files or documents with special confidentiality needs?
- Who may see which sources? An assistant should only use content the asking person can access themselves.
- Which actions run freely, which only after a prompt and which not at all? Anything that acts outward matters most.
- Who reviews an answer that leaves the company before it is sent?
- Which contracts are needed? If an external service is used, a data processing agreement under Art. 28 GDPR usually belongs to it.
- How is the works council involved? Technical systems that can be used to monitor the behaviour or performance of employees may be subject to co-determination under section 87(1) no. 6 of the Works Constitution Act, and an assistant that keeps usage records can fall into that category.
This is not legal advice. Clarify your individual case with your legal counsel and your data protection officer. In practice: involve the works council in week 1, not in week 6. A pilot that has to be halted midway costs more trust than any early question.
The permissions for actions should also be written down as a rule and be visible in the tool. In TheroAI it looks like this: for each tool you can choose Allow, Ask or Block, and writing actions are marked. Whatever the product, the question is the same: can an administrator see at a glance what the assistant is allowed to do?
At the end of week three comes gate 1: is the use case chosen, are the baseline values collected, do the rules exist, and is the works council informed? If so, the pilot starts. If not, the gaps are closed and the start moves. That is not failure but the purpose of a gate.
Phase 2: Pilot (Weeks 4 to 9)
| Week | Focus | Result at the end of the week |
|---|---|---|
| 4 | Onboarding the six people | All have worked on real tasks with the assistant |
| 5 | First routine, collect hurdles | List of hurdles, first adjustments to instructions and templates |
| 6 | Interim review | Conversation with those who barely use the assistant |
| 7 | Measurement begins | Usage is recorded, sampling of answers runs |
| 8 | Measurement, second time check | Times for the ten tasks from weeks 2 and 3 recorded again |
| 9 | Evaluation | Traffic light filled in, gate 2 with a minuted decision |
The Rhythm: 30 Minutes Once a Week
The pilot lives on a short, fixed review. The group meets once a week for 30 minutes and answers three questions: What helped? What got in the way? What did you not use although you could have? The third question is the most revealing. It shows where habit is stronger than the tool, and whether that comes from missing onboarding, distrust of the answers or the wrong task.
Expect the first week or two to be bumpy. The pilot group learns to ask questions in a way that produces usable answers, and it learns to read answers critically. A library of reviewed templates for typical requests helps because it makes good wording shareable. Keep adjustments in writing so it stays clear later what changed between week 4 and week 9.
Measuring Without Monitoring
Weeks 7 to 9 are for measuring. Three things are enough: How often do the six people use the assistant? How good are the answers? How long do the ten tasks take compared with the baseline?
For quality, the project owner and a subject expert draw 20 answers in total from the pilot conversations. What matters is that they are drawn at random and not hand-picked. Each answer is checked against the source it names, and counts as wrong if its content is incorrect or it has no evidence. For time, checking the answer is part of the task. An assistant that replies in two minutes but whose result needs ten minutes of verification has not saved time.
Whether usage comes from your tool’s logs or from a short poll of the group depends on what your tool provides. As soon as usage is recorded, it belongs in the conversation with the works council from week 3. Agree that what is evaluated is what the group does as a whole, not how diligent individuals are. The figure “five of six use the assistant regularly” is entirely sufficient for the decision. Regular here means use on at least three working days per week.
Week 9: Traffic Light and Gate 2
In week 9 the measurements go into a simple traffic light, which we describe in the section on stopping criteria. Gate 2 is the most important decision in the plan: this is where you stop, improve or expand. Plan two hours for it and invite the sponsor in person.
Phase 3: Expansion (Weeks 10 to 13)
| Week | Focus | Result at the end of the week |
|---|---|---|
| 10 | Onboard a second group | Members of the first group accompany the second |
| 11 | Organize operations | Responsibilities, a contact point for questions, a way to make changes |
| 12 | Check access and cost | Permissions spot-checked, effort and licence costs compiled |
| 13 | Closing report | Measurements of both groups, recommendation, gate 3 |
Expansion does not mean everyone starts at once. It means a second group, a second round of measurement, and a check on whether the first group’s experience carries over. Members of the pilot group accompany their new colleagues. That shortens onboarding and gives the pilot people the role of multipliers, which they have earned.
Operations Is a Task of Its Own
The expansion phase does not end with more users but with a responsibility. Who answers questions? Who maintains the templates and instructions? Who spot-checks whether permissions are still right? Who decides on new use cases? If after week 13 nobody is entered by name, operations become accidental and usage falls apart.
A tool should help keep processes traceable. If part of the work runs as an automated flow, for example a customer reply that is only sent after approval, a log of the runs helps. In TheroAI the audit log for flows shows the time, status and duration of each run, filterable and exportable as CSV. More on flows with an approval step is on the workflows page.
At the end of week 13 comes the closing report. It contains the baseline values, the measurements of both groups, the sampling results, the open risks, the effort and the recommendation. It does not need to be longer than two pages.
The Gates: Three Decisions Instead of One
The most common design flaw in AI projects is a single decision point at the end, where the project can then only be praised. Anyone who has already invested money, time and reputation rarely decides against themselves. Three gates, each of which also allows a stop, spread the decision over moments when little has been lost.
At each gate four kinds of input come together: usage, the quality of answers, the group’s feedback and the open risks. The round consists of the sponsor, the project owner, the business unit, and IT and data protection. It has three possible outcomes, and all three carry equal weight.
| Gate | Timing | Question | Possible decisions |
|---|---|---|---|
| Gate 1 | Day 21 | Are we ready for the pilot? | Approve the pilot start, improve first or call it off |
| Gate 2 | Day 63 | Does the pilot support expansion? | Continue, adjust or stop |
| Gate 3 | Day 91 | Do we take over operations? | Take over operations, improve first or end |
The decision is minuted with reasons, however it turns out. A stop without reasons leaves only the feeling that it did not work. A stop with reasons leaves knowledge that is useful next time.
Typical Pitfalls
Planning projects like this reveals recurring mistakes. The following table is a compilation from our point of view, not a statistic.
| Pitfall | How you notice | Remedy |
|---|---|---|
| No clear goal | Everyone explains differently what the assistant is for | One goal sentence in week 1, signed by the sponsor |
| Too many use cases | The group spreads itself thin, nothing becomes solid | One case, one team, one review rule |
| No baseline | After the pilot people say “feels faster” | Time measurement in weeks 2 and 3 |
| Pilot only for enthusiasts | Everyone in the pilot is happy, nobody outside uses it | Mix the group, invite sceptics too |
| Works council and data protection too late | The pilot is halted in week 6 | Involvement from week 1 |
| Answers are not checked | Mistakes surface only at the customer | Sampling, source requirement, approval for anything with external effect |
| Ownership ends with the pilot | After week 13 nobody feels responsible | Name operational ownership in phase 3 |
Several of these pitfalls are linked. Anyone who formulates no goal in week 1 cannot justify a use case in week 2 and cannot evaluate anything in week 9. The plan is therefore built as a chain: each week delivers what the next one needs.
One pitfall deserves its own paragraph because it cannot be dealt with in a table row: the answers themselves. Language models phrase things convincingly even when they are wrong. A tool that names its sources makes checking easier but does not replace it. For anything that goes outside the company, human approval before sending is the simplest and most reliable rule.
Stopping Criteria: When Quitting Is the Right Call
A plan that does not provide for quitting will not quit when it matters. So set the criteria before the pilot starts, while nobody has a stake in a particular result. The traffic light below shows four signals with example thresholds. They are not industry values. They show the method, and you should set your own thresholds.
Here is how to read the traffic light:
- Usage: With six pilot people, green means five to six use the assistant regularly, yellow means three to four, and red means zero to two. A technology nobody uses voluntarily has a problem that more training rarely solves.
- Quality: Of 20 checked answers, at most one may be wrong or unsupported for the field to stay green. Two to three are yellow, four or more red.
- Time per task: Green means at least 20 percent shorter than the baseline, yellow means less than 20 percent shorter or equal, red means longer than before. Checking time is included.
- Open risks: Green means settled and documented points on data protection and the works council, yellow means open points with a fixed date to settle them, red means open points with no solution in sight.
The rule for evaluation: no red and at most one yellow field means continue. One red field or several yellow ones mean adjust. Two or more red fields mean stop. An exception to the rule is possible, but only in writing and with reasons from the sponsor.
Regardless of the traffic light, there are grounds for an immediate stop: an answer containing confidential content for the wrong person, a breach of the rules from week 3, or an unresolved legal question that cannot wait. Then the pilot is halted before the next traffic light is due.
A Stop Is a Result
A stop is not a failure as long as it leaves something behind. Record which threshold was missed, what you learn from it, and whether a different use case, a different time or a different tool justifies a new test. A company that can say after 13 weeks “For this task it is not worth it, for these reasons” has gained more than one that says after 13 weeks “We rolled it out”.
We would be glad if your pilot ended with green fields, and as a vendor we have an obvious interest in that. That is exactly why we consider the stopping criteria the most important part of the plan: a pilot that is not allowed to fail says nothing about the benefit.
What You Can Take Away
If you are thinking about an AI rollout in your team, these questions help, whichever tool you choose:
- 1.Can you write the goal in one sentence, with a team, a task, a date and a yardstick?
- 2.Do you have a sponsor who releases time and accepts an uncomfortable result?
- 3.Have you chosen a use case that occurs often, does little harm when it goes wrong and can be checked against a source?
- 4.Have you collected baseline values before the pilot begins?
- 5.Are rules for data, access, approvals and the involvement of the works council in writing?
- 6.Are the stopping criteria set before anyone has a stake in the result?
- 7.Is it settled who owns operations after week 13, should it continue?
Anyone who answers these seven questions with yes does not yet have a successful AI rollout. They have one that can be checked. After 90 days that is worth more than a promise on day one.
If you would like to see what an assistant with sources, permissions and approvals looks like in daily work before you put your group together, you can take a look at the demo at your own pace. A roadmap works regardless of which tool ends up in your pilot.
See Thero live
Book a short demo. You talk directly to the founding team.