#Observability
#AIOps
#OperationalAI
#OperationalEngineering
From Mean Time to Innocence to Mean Cost of Innocence


Nick Randall
September 7• 3 Min read

What does it cost to reach a defensible operational conclusion?
TL;DR
Operations teams have spent decades trying to reduce Mean Time to Innocence — how quickly they can establish whether their part of a system is responsible for a problem.
As more operational reasoning moves from human experts into automation and AI, perhaps we should also measure Mean Cost of Innocence — what it costs to reach that defensible operational conclusion.
For much of my career in networking, a significant part of the job has been proving that the network is innocent — or finding out that it isn't.
We used to talk about Mean Time to Innocence (MTTI): how quickly could we get from an alarm or reported problem to enough evidence to establish whether the network was responsible?
Traditional monitoring was actually quite efficient at dividing the work between people and machines.
We collected telemetry, configured thresholds and rules, and generated events. The machine identified something worth investigating; an experienced engineer supplied the reasoning.
That engineer brought years of knowledge to the problem. They knew which measurements mattered, which other systems might be involved, what sequence of checks to make, and what constituted sufficient evidence to reach a conclusion.
Increasingly, we want machines to perform some of that reasoning too.
That changes the economics.
One approach is to give an AI access to enormous quantities of operational telemetry and ask it to work things out. But if the knowledge required to interpret that telemetry already exists, why repeatedly pay a machine to rediscover it?
The knowledge may sit with an engineer inside the organisation. But it might equally reside with an equipment manufacturer, software vendor, service provider or another specialist.
The interesting question is therefore how we capture and operationalise more of the specialist knowledge that already exists before repeatedly paying people or machines to reconstruct it.
Data isn't the same thing as information
An operational system might generate millions of measurements.
That doesn't mean every measurement provides millions of pieces of new operational information.
An experienced engineer understands this intuitively.
They know that a changing measurement may be insignificant. They know that several measurements considered together may indicate a particular operational state. They know which observations provide useful evidence for a diagnostic question and which are largely irrelevant.
In other words, the expert doesn't simply consume data.
They interpret it.
If we can capture appropriate parts of that interpretation, we can start transforming raw operational data into information that is better suited to whoever — or whatever — needs to consume it next.
Sometimes that consumer will still be a human.
Increasingly, it may be automation or AI.
The cheapest reasoning isn't always AI — or a rule
Some operational questions can be answered with a simple threshold.
Others are better addressed with deterministic analytics or statistical techniques.
Some genuinely benefit from AI inference.
And some still warrant the attention of an experienced human.
There isn't a universal hierarchy.
A few pence of AI inference may be considerably more economical than engineering and maintaining a bespoke analytical workflow. Equally, repeatedly sending huge quantities of operational data to a general-purpose reasoning engine may make little sense when established domain knowledge can first transform that data into a much more useful representation.
The objective shouldn't therefore be to minimise telemetry, tokens, AI or human involvement for their own sake.
It should be to ask:
What combination of data, transformation, context and reasoning produces the required operational outcome most economically?
That's a system-engineering problem as much as an AI problem.
The cost isn't only compute
There is another cost that is easy to overlook.
Operational knowledge often has to cross boundaries before it can be used.
A network specialist may need an observability engineer to implement their reasoning. An enterprise might depend on an equipment vendor to interpret a particular set of measurements. A service desk might escalate to several teams before reaching somebody with enough knowledge to diagnose a problem.
Each hand-off can add time, engineering effort and loss of context.
So the cost of operational reasoning isn't simply the cost of storing telemetry or running an AI model.
It can include:
telemetry ingestion and storage
queries and compute
engineering time
DevOps and platform resources
external specialist involvement
organisational escalation
AI tokens and inference
and, of course, elapsed time
As more operational decisions move towards automation and AI, making specialist knowledge reusable rather than repeatedly accessible only through scarce people and systems becomes part of the economics too.
What does it cost you to reach an operational answer?
A LeanMetrics Assessment examines a real operational workflow — its telemetry, expertise, tooling and reasoning — to identify where time and processing cost could be reduced.
Explore a LeanMetrics Assessment
From Mean Time to Mean Cost
That brings me back to Mean Time to Innocence.
MTTI remains useful:
How quickly can we reach enough operational understanding to establish whether a particular domain is responsible for the problem?
But perhaps it now needs a companion.
Mean Cost of Innocence (MCOI)
What resources did we consume reaching that defensible operational conclusion?
The two measures are related, but they're not the same.
We could make diagnosis faster by throwing enormous amounts of compute, telemetry and AI inference at every problem.
We could make individual technology costs lower while consuming hours of scarce engineering time.
We could reduce telemetry aggressively and discover that we've removed information somebody later needs to reach the correct conclusion.
None of those necessarily represents better operational engineering.
The objective is to reach a sufficiently confident operational conclusion using an appropriate combination of resources.
We've spent years trying to reduce Mean Time to Innocence.
As we automate more of the reasoning traditionally performed by operational experts, perhaps we should pay equal attention to Mean Cost of Innocence.
Resources
Copyright NetMinded, a trading name of SeeThru Networks ©
