← Compendium

Getting KPIs right

You don’t set up a KPI because you can measure something, but because you need visibility into something specific: a number that should improve, or a guard that warns you when a change elsewhere makes something good worse. Everything else is a vanity metric. The purpose dictates the form, a corridor with a target, a risk threshold and an opportunity threshold instead of a bare number, and running a product takes just three of them: allocation, forecast accuracy and efficiency.

Many dashboards are full of numbers that lead nobody to a decision. They are shown because they can be shown, but none of them says whether something is going well right now or whether someone ought to step in. That is the core of the problem: things are measured because they can be, not because a particular question needs answering.

Before the first metric, it is therefore worth looking not at formulas but at the purpose. What exactly do we want to see, why do we want to see it, and which decision would a change in this number trigger? A metric that has no answer to this question is a vanity metric: it looks nice and decides nothing. Only once the purpose is clear does a metric make sense, and only then is it worth talking about its form.

What a KPI is for

At its core, a KPI has one of two roles. In the first, it is a driver: there is a number that should get better, and the KPI makes progress visible so you can tell whether your measures are working. In the second role, it is a guard. Then it is not about improving the number, but about making sure it does not tip over while you deliberately change something elsewhere.

The guard role in particular is often overlooked, although in everyday work it is the more valuable one. Anyone who speeds up delivery wants to be sure that the error rate or the costs do not quietly run away in the process. Anyone who cuts costs wants to notice when reliability suffers as a result. A guard KPI is thus the guardrail that protects a planned change in one place by raising the alarm in another before the damage becomes visible.

Both roles lead to the same consequence: a metric must not only deliver a value, but be able to say whether that value is all right. A driver needs a target to work towards, a guard a limit beyond which it warns. A bare number can do neither, and that is exactly where most dashboards fail.

A KPI is a corridor

For a metric to drive or warn, it needs a range rather than a single point. A useful KPI consists of a target value, a risk threshold and an opportunity threshold. The target value says where you want to be. The risk threshold is the guard’s limit, beyond which you have to act. The opportunity threshold shows from when something is going better than planned, so a closer look is worthwhile to learn from it.

A KPI is a corridor, not a point Opportunity better than planned Target corridor expected range around the target value Risk worse than planned opportunity threshold risk threshold target value
A KPI spans a range: the target corridor around the target value (green), the risk threshold beyond which the guard warns, and the opportunity threshold beyond which a closer look is worthwhile. A bare number without these thresholds does not say which of the three ranges it is in.

Two more things belong to a useful KPI. It rests on a deterministic query that always delivers the same value from the same data, instead of an estimate that comes out differently every time. And it is versioned, because expectations change over time and you need to be able to trace which threshold applied from when and why. Only then does a metric become a tool you align decisions with, instead of just looking at numbers.

Three KPIs for cloud costs

So much for the general mechanics. Purpose, driver, guard and corridor apply to every metric. Now it gets concrete, with the question to which we apply these mechanics most often: cloud costs. For a product company they are an ongoing, moving item, and out of the many possible metrics, three have proven themselves here that give the greatest insight with minimal effort. They answer three questions: do we know what our costs are for, can we foresee them, and do we make use of what we provision? Each is thought of as a corridor, so it can drive or warn instead of just showing a number.

Three KPIs for cloud costs 01 Allocation Do we know what we are paying for? 75% of costs allocated 02 Forecast accuracy Can we foresee them? 93.75% forecast on target 03 Efficiency Do we make use of what we provision? 69% used efficiently
The three cloud cost KPIs at a glance: allocation asks what we are paying for, forecast accuracy whether we foresee it, and efficiency whether we make use of what we provision. The percentages are the example values from the following paragraphs; each is meant as a corridor with a target and a risk threshold, not as a bare number.

The first is allocation. It answers the simplest and at the same time most uncomfortable question, namely whether every cost item can be assigned to a product, a service or an area. In the cloud, orphaned resources quickly appear that run and cost money, but for which nobody feels responsible. The allocation KPI measures what share of the costs has a clear assignment. If a month’s total cloud costs come to €8,000 and €6,000 of that can be assigned to specific products, allocation is 75 per cent, while the remaining 25 per cent stays in the dark. As a corridor, the target value could be 90 per cent and the risk threshold 70 per cent, below which too much remains unassigned.

The second is forecast accuracy. Cloud costs are not fixed, because they fluctuate with usage and change with every new service, and this metric measures how well you foresee that movement. If an additional €8,000 is expected for a new project, but the actual costs at the end of the month come to €7,500, the absolute error is €500. Accuracy is calculated by relating this error to the forecast and subtracting it from the whole, that is (1 minus 500 divided by 8,000) times 100, which equals 93.75 per cent. What matters is that high accuracy counts in both directions, because both clearly overshooting and clearly undershooting show that you do not yet have your own consumption under control.

The third is efficiency. The strength of the cloud is its dynamic scalability, but that raises the question of whether what is provisioned is actually used, or whether unused capacity is quietly tying up capital. The pragmatic way in is not to measure every resource exactly, but to split the costs into two groups: managed services are assumed for now to have an efficiency of 100 per cent, compute resources 50 per cent. If, out of total costs of €8,000, about €3,000 go to managed services and €5,000 to compute, the efficient costs are 3,000 plus half of 5,000, that is €5,500, which equals roughly 69 per cent. These starting values are deliberately rough and probably not correct, because their purpose is to challenge the status quo and invite the teams to provide more accurate values.

The pragmatic way in

These three metrics can be brought to a useful level within a few weeks, without waiting for a perfect system. For allocation, you capture the cloud costs and define which product or area is responsible for which accounts and resources. For the forecast, a weekly estimate compared with actual developments is enough at the start. For efficiency, you set the rough starting assumptions and refine them step by step.

What matters is the order and the attitude behind it. You deliberately start with a “good enough” state and build on it, instead of getting lost in details such as choosing between individual instance types before the basic structure is in place. Once the three KPIs are set up as corridors, they are best carried by a deterministic query that delivers the value reproducibly, so the recurring look at the numbers turns into real steering.

There are several options for the place where these corridors come together, two of which are particularly worth a look. The first is a ready-made tool: we ourselves use our Startup Business Cockpit, where you enter an indicator while the calculation and chart are created automatically, and where you define which metric is shown to all employees at which level of detail, so the target corridor stays visible to those who need it without every number being open to everyone. The second is building it yourself on a queryable data foundation that a language model queries in natural language, as we describe in the article Talking to your own numbers. Both lead to the same goal: a corridor that is calculated reproducibly and visible to the right people.

What this means for you

A metric does not become useful by existing, but by serving a clear purpose: making visible something that should change, or safeguarding something that has to stay good. The form follows from this purpose, namely a corridor with target, risk and opportunity that comes from a reliable query. Running a product does not take a KPI zoo, but a few metrics that drive or warn. Whether it is about the first metrics or about finally making existing reporting fit for steering, we are happy to sort it out together with you.

Frequently asked questions

What is a KPI in this sense?

Not a single target value, but a corridor: a target value plus a risk threshold and an opportunity threshold. Only this range turns a number into a statement you can align decisions with, whether as a driver or as a guard.

Why set up a KPI at all?

Because you need visibility into something specific: a number that should improve, or a guard that warns when a change elsewhere makes something good worse. Measuring without this intent only produces a vanity metric that looks nice and decides nothing.

Why only three KPIs?

Because a metric can be invented for every situation, and a KPI zoo smothers exactly the steering it is supposed to enable. Allocation, forecast accuracy and efficiency answer the decisive questions with minimal effort: do we know what we are paying for, can we foresee it, and do we make use of what we provision?

Why a deterministic query instead of a dashboard?

Because a KPI has to be reproducible: the same data must always give the same value, otherwise you argue about the number instead of the decision. A deterministic query, versioned and traceable, delivers exactly that, whereas an isolated dashboard tile only shows without putting anything into context.

Related topics