Imagine a quarterly review. In spring a company introduced an AI assistant along with three workflows: preparing customer replies, summarizing meetings and drafting a weekly report. The finance lead asks a simple question: “What has this brought us?” The project lead answers honestly: “People like it, and I think we save a lot of time.”
The scenario is a thought experiment, not a customer case. The answer is still typical, and it is not wrong. It is just a claim. Someone who hears a claim has to believe it or disagree. Someone who sees evidence can check it, recalculate it and ask follow-up questions. Between the two there is no huge effort, but there is a series of decisions that have to be made before the start and not in the quarterly review.
This article describes five metrics for an enterprise AI: usage, success rate of workflows, estimated time saved, quality and time to answer. It shows where you find them in the TheroAI insights, what they are good for and, just as important, what they are not good for. It includes a worked example that is clearly labeled as an example.
Our thesis: value can be proven if you keep three things apart, namely what is measured, what is estimated and what people judge. And if you record the before value before the AI arrives in daily work. Whoever does that can defend even a modest number with confidence. Whoever does not can defend even a large number only on faith.
Why “People Like It” Is Not Enough
Approval is valuable. Without it no tool gets used, and an unused tool delivers nothing. But approval answers a different question than the one finance asks. It says that something is liked. It does not say that something works. Three reasons why the two drift apart in practice:
Novelty is not a permanent state. In the first weeks many people try out what is possible. Usage numbers are high, the mood is good. Whether that turns into a habit only shows after months, and only a number over time can show it.
Perceived time saved is unreliable. Someone who gets a draft in seconds remembers the seconds. The minutes spent afterwards checking and adjusting fade faster. That is not dishonesty, it is perception. This is why you need a comparison with something that was recorded before the start.
Without a before value there is no after. If nobody knows how long a customer reply took before, nobody can say whether it takes less time now. That gap cannot be closed in hindsight. You can ask people how it used to be, but that is memory again.
There is also a practical reason. A budget that is up for renewal needs a justification that someone who has never opened the assistant can understand. A number with a source is such a justification. A mood is not.
Five Metrics That Belong Together
There are countless metrics you could collect. We think five are enough as long as you read them together. Each answers exactly one question, and the questions build on each other: Is the AI being used? Does it run through? Does the result help? Does that free up time? And does the answer reach the other side faster?
The graphic also shows what most reports leave out: not all numbers carry the same kind of truth. Some things are counted, some are calculated, some are a person's judgment. The table below sorts the five metrics and names the most common trap for each.
| Metric | Question behind it | Where you find it | Kind of number | Most common trap |
|---|---|---|---|---|
| Usage | Is the AI needed in daily work? | Dashboard: runs, total processes, range of 7, 30 or 90 days | measured | Early curiosity is mistaken for habit |
| Success rate | Does the workflow run through? | Dashboard: success rate, audit log: status | measured | Finished does not mean correct |
| Quality | Does the result help? | Feedback: helpfulness rate, rating | human judgment | Few responses, skewed selection |
| Time saved | Does working time get freed up? | Dashboard: estimated time saved | estimated | An assumption looks like a measurement |
| Time to answer | Is it faster? | Audit log: duration, average duration | measured (run time) | Run time is not waiting time |
In TheroAI these numbers sit in one place: the insights, with the four tabs Dashboard, Audit-Log, Feedback and Muster (patterns). The order of the following sections follows the order of the questions.
Usage: Who Asks, How Often, and What It Does Not Tell You
Usage is the simplest metric and the one most often misread. The dashboard shows the number of runs, the total number of processes and an overview by department. You can choose a range of 7, 30 or 90 days. There is a PDF export for the report.
What you should read from it is rarely the total. Three questions say more:
- What does the curve look like? Does the number of runs rise over 30 and 90 days, stay flat or fall back after an initial peak?
- How broad is the usage? Do all the workflows that were set up actually run, or only one? A single workflow that produces 90 percent of all runs is a finding, not a success.
- Where does it sit? The department overview shows in which groups the workflows really run and which workflow is executed most often there.
One limit belongs here. The evaluation counts runs of agents and workflows. How often people work with the assistant in chat is a different kind of usage. If you want to assess it, you need a separate measurement, for example a short survey after the pilot or a sample of conversations in the team. A number you do not have should not be derived from another one.
And one more warning against the reverse conclusion: many runs are not automatically much value. A workflow can run a hundred times a day and deliver a result every time that nobody reads. Usage is the ticket of admission for the other metrics, not their replacement.
Success Rate: Finished Is Not Correct
The second question is whether a workflow runs through. The dashboard reports the success rate for this, and in the audit log every run has a status: Erfolgreich (successful), Fehlgeschlagen (failed) or Ausstehend (pending). With “Statistik anzeigen” (show statistics), administrators display four figures above the table: total, successful, failed and the average duration.
The success rate is a technical number. It says that a run reached the end. It does not say that the result was right. A workflow that reliably delivers a summary with the wrong emphasis has a success rate of 100 percent and a value of zero. Conversely, a low rate is not always a quality problem: if a connection goes down or a permission is missing, the run fails even though the workflow itself is good.
That is why the success rate should never be shown alone. Two things belong with it: a look at the failed runs before you dismiss them as noise, and the quality number from the feedback. How the two interact is shown below in a small matrix.
Estimated Time Saved: Why “Estimated” Is the Honest Word
Time saved is the number everybody wants to see, and the one where the most corners get cut, often unintentionally. In the insights it is deliberately called “Geschätzte Zeitersparnis” (estimated time saved), and that adjective is not habitual caution but the correct description.
On the calculation: if a comparison time is stored for a workflow, meaning the minutes the work takes by hand per case, the evaluation multiplies the number of runs by that comparison time. How and by whom that time is stored for your workflows is best settled during setup. Without a stored comparison time it uses a flat assumption: three times the average run time, at least five minutes per run. This flat rate is a placeholder. It is better than an empty column, but it is not a measurement. For the figures without a comparison time that means: order of magnitude, not proof.
Why “estimated” remains the right word even with a carefully recorded comparison time comes down to three reasons:
- 1.The comparison case does not exist. Nobody can observe how long the same task would have taken by hand while the AI was doing it. You can only carry over a comparison time from earlier cases.
- 2.The checking time is missing in the simple calculation. Runs times comparison time assumes that the AI takes over the entire job. In reality a human reads, corrects and approves the result. That time has to be deducted.
- 3.Freed time is not the same as used time. If ten minutes are freed up, it is open whether they turn into ten minutes of productive work. It stays capacity, not money in the bank.
A clean reporting practice therefore separates three things that often blur together in presentations: the observed time (measured before and after), the capacity value (freed hours, calculated) and the financial saving (only if costs actually fall). Whoever derives the third number from the first without mentioning the second claims more than can be proven.
If you want to run the calculation with your own assumptions, the ROI calculator helps. It does not replace a measurement, but it forces you to write the assumptions down.
The Baseline: Record the Before Value Before You Start
The baseline is the measured value before the start. It is the least conspicuous and most important ingredient of the whole proof, because it is the only one that cannot be caught up on. Half a year after the start nobody can measure how things were before.
The effort is manageable. Pick two or three workflows for the beginning, not twenty. In the four weeks before the start, record on real cases what you want to compare later:
| Quantity | How to record it | What it is for later |
|---|---|---|
| Handling time per case | Ten to twenty real cases with a stopwatch or a self-note, take the median | Comparison time per workflow |
| Time to answer | Timestamps of receipt and reply from mailbox or ticket system | Comparison of total duration |
| Cases per week | Count over four weeks | Putting usage in context |
| Rework and follow-up questions | Small sample: how often does a correction or follow-up question come in? | Comparison of quality |
The median is better than the average here, because single outliers, such as a special case that took half a day, distort the average. The number should represent the typical case.
Two practical notes. First: measure the workflow, not the person. The question is how long a customer reply takes, not who is slower. Second: talk to the people involved beforehand. A stopwatch without an explanation breeds mistrust, and mistrust colors the measured values.
The timeline is a proposal, not a rule. What matters is the order: before value first, then the start, then a settling-in phase that is not counted because everybody is still learning in the first weeks, and only then the measurement phase. The numbers from the settling-in phase do not belong in the report, otherwise you penalize the project for something that is normal.
Quality: What Feedback Says and What It Does Not
Whether a result helps can only be judged by people. That is why TheroAI has feedback: after a run the user is asked whether the result was helpful, and the answer flows into the evaluation. In the Feedback tab you see the helpfulness rate, an average rating, the total number of responses and those of the last seven days, as well as the best-rated workflows and the ones with the most room for improvement. When there are too few responses, the evaluation itself points out that there is not yet enough data for a comparison. That is the right attitude.
Still, this number is the easiest one to overrate. Four traps:
- Low response rate. If only 60 of 180 runs have a response, the statement covers a third of the cases. Always state both numbers.
- Skewed selection. Feedback comes mostly from people who were very annoyed or very pleased. The middle stays silent.
- Helpful is not correct. A fluently worded result feels helpful even if a figure in it is wrong. Feedback does not replace checking against the sources.
- Rating depends on expectation. Whoever expects a lot gives worse grades. A falling rate can therefore also mean that expectations have risen.
A small sample with a professional check works as a counterweight: ten results a month, drawn at random, are checked by a subject matter expert against the cited sources. That costs an hour and complements the feedback with a judgment that does not depend on the mood of the day.
It is especially revealing to read success rate and feedback together. The matrix below shows what the four combinations mean.
The most dangerous field is the one with a high success rate and poor feedback. It looks green in the dashboard. Whoever reports only the success rate overlooks it. Typical causes are an instruction in the agent that is too vague, missing or outdated sources, or a task that is not suited to an automatic workflow.
Time to Answer: Two Clocks
The fifth metric sounds simple and is most often confused with something else. There are two clocks. The first measures how long a run works. In the audit log it sits in the column Dauer (duration), in the product capture above for example as 2.8 seconds or 1.6 minutes, and the statistics show the average duration. The second measures how long a person waits for the answer. Between the two lies everything people do and leave undone: a request sits in the mailbox, a draft waits for review, an approval is outstanding.
For workflows with an approval step the difference is large. An approval step has a deadline, in the example of the product captures it is 72 hours. The actual work of the run can be done in a few seconds, yet the answer reaches the customer only once someone has approved it. Whether the duration shown in the audit log includes the wait for an approval is something to check on your own runs before you report it as run time. That is not a flaw but the intention of a four-eyes principle. It does mean, though: whoever sells run time as time to answer is flattering the numbers.
So measure the second clock where the person feels it: at the timestamp of the incoming request and at the timestamp of the sent reply, from the mailbox or the ticket system. You made the same measurement in the baseline. Here too, median beats average.
The graphic shows a worked example with round assumptions, not a measurement. But it makes visible what often happens in practice: the run time of the AI is a tiny part of the total time. The bigger lever sits in the waiting. Whoever wants to shorten the time to answer therefore looks not only at the machine but also at who is notified when and how the review is organized.
From Number to Improvement: Patterns, Audit Log and Report
Metrics are not an end in themselves. They are only worth something when they lead to decisions: expand a workflow, sharpen an instruction, repair a connection, end a pilot. The insights offer two more building blocks for this.
The Muster tab shows “Gelernte Muster” (learned patterns). With “Analysieren” (analyze), TheroAI evaluates runs and suggests patterns, such as a preferred tool choice, a recurring decision rule or a workflow pattern. Each pattern states how many runs it is based on, and a person confirms or rejects it. That is the right division of labor: the system finds candidates, the business side decides. A pattern is a hint from the runs, not proof, and the number of underlying runs says how much weight it deserves.
The second building block is the audit log as evidence. If someone asks in the quarterly review how a number came about, the list of runs can be filtered and exported as CSV. The number then has a source. That is what separates evidence from a slide.
A monthly report does not have to be long for this. One page is enough if it contains these five things:
- the five metrics with the period and the comparison to the baseline
- for each number its kind: measured, estimated or judged
- the number of responses in relation to the number of runs
- the assumptions behind the estimated time saved, listed one by one
- one decision or question for the next month
The last point is the most important. A report that only describes gets read and forgotten. A report that prepares a decision gets used.
A Worked Example, Step by Step
The following calculation is a worked example with round, invented assumptions. It shows the path from the naive to the reliable number. The workflow is fictional: preparing a customer reply and having it approved, modeled on our demo materials with the invented error E-217 on the filler FL-200.
| Quantity | Value | Kind |
|---|---|---|
| Runs per month | 200 | assumption |
| of which successful | 180 (success rate 90 percent) | assumption |
| Manual work per request before | 20 minutes (median from the baseline) | assumption |
| Checking and adjusting per request with AI | 8 minutes | assumption |
| Feedback | 60 submitted, of which 48 helpful (80 percent) | assumption |
| Time to answer | before 6 hours, after 3 hours (median) | assumption |
Now the calculation in three steps:
| Step | Calculation | Result |
|---|---|---|
| 1. Naive: all runs times comparison time | 200 times 20 minutes = 4,000 minutes | about 67 hours |
| 2. Subtract failed runs | 20 times 20 minutes = 400 minutes | minus about 7 hours, 60 hours remain |
| 3. Subtract checking time | 180 times 8 minutes = 1,440 minutes | minus 24 hours, 36 hours remain |
67 hours become 36. That is still a solid result, and it is one that withstands a follow-up question. The 67 hours would not have held up at the first critical question: where is the checking time? What about the runs that failed?
To put the other values in context: with 60 responses for 180 successful runs, the statement “80 percent helpful” rests on a third of the cases. That is a usable signal but not a blank check. The time to answer is cut in half in this example, from 6 to 3 hours, and most of it is still waiting time. And the 36 hours are freed capacity, not saved money. Whether that turns into more output, shorter overtime or nothing at all is a decision, not a measurement.
Data Protection and Co-Determination: Measuring Without Monitoring
Whoever measures usage creates data about behavior. That is unavoidable and fine as long as it is handled consciously. The audit log has a Person column. The dashboard groups by department. For proving the value, the level of the workflow and the department is almost always enough. You do not need a ranking of individual employees by usage for this, and it would harm the effort more than help it.
In practice that means: define the purpose of the evaluation beforehand and write it down. Clarify who may see which evaluation. And involve the works council early. Under section 87 (1) no. 6 of the German Works Constitution Act, technical systems suited to monitoring the behavior or performance of employees are subject to co-determination. Whether and how that applies to your situation is a matter for the individual case. This is not legal advice; clarify the individual case with your legal counsel.
On where the data lives: data is kept in Germany and AI processing takes place in the EU. You will find more on the security page. What matters for the metrics is a different point: you measure workflows, not people, and you tell the people involved before you begin.
What You Can Take Away
If you have to prove the value of an enterprise AI, these steps and questions help:
- 1.Before the start, pick two or three workflows and measure the before value for four weeks: handling time, time to answer, cases per week and rework, each as a median.
- 2.Plan a settling-in phase of about four weeks that does not count toward the assessment.
- 3.For every number, check whether it is measured, estimated or judged, and write it next to it.
- 4.Never show the success rate without quality. Ask: what is the response rate, and who checks samples against the sources?
- 5.Calculate time saved net: deduct failed runs and checking time, list the assumptions one by one and do not book capacity as money.
- 6.Measure the time to answer where the person waits, not at the run time of the system.
- 7.Clarify purpose, visibility and co-determination before the first evaluation is produced.
Could someone who has never used the assistant recalculate from your report how the numbers came about?
If the answer is yes, you have evidence. If not, you have a claim, even if it is true. Our position: an enterprise AI does not need to make its value bigger than it is. It needs to show it in a way that can be believed, and for that five numbers, a before value and the willingness to write out the word “estimated” are enough.
If you would like to see how the insights with dashboard, audit log, feedback and patterns look, we are happy to show them in a demo, without obligation and at your pace.
See Thero live
Book a short demo. You talk directly to the founding team.