Articles

Data Sovereignty in the Age of Autonomous Observability

Rajith Attapattu

Rajith Attapattu

September 27, 2026 · 8 min read

Originally published on rajithattapattu.com

ObservabilityData SovereigntySRE AgentIncident Response
Data Sovereignty in the Age of Autonomous Observability: an agent reasoning next to the data inside your environment, sending only findings and evidence to the control plane

In short

Building an autonomous SRE Agent showed that more raw context produces better root-cause analysis than pre-distilled signals, which conflicts with observability pricing that pushes teams to keep less data. Instead of centralizing more telemetry, move the agent's reasoning to where the data lives and send only findings, evidence and summaries to a central control plane.

Before we started working on AI-driven observability using autonomous SRE Agents, a lot of our focus was on correlating raw telemetry to extract high-value signals.

It made sense. Humans take time to process raw telemetry and connect the dots. If we could do some of that work ahead of time and turn large volumes of metrics, logs, traces and events into higher-value signals, engineers could get to the problem faster.

So when we started building our SRE Agent, we initially used those high-value signals as the input.

The thinking was that we had already distilled the data and extracted what really mattered. The agent could take those signals, connect the dots, reason about what was happening and get to the root cause much faster.

We were wrong.

With highly distilled signals, the agent lacked the context required for deeper reasoning. When we started giving it access to more of the underlying data, we noticed the quality of its reasoning improved significantly.

Think about how an experienced engineer troubleshoots an incident.

They might look at metrics, logs, traces, Kubernetes events, configuration, infrastructure state and historical information such as previous incidents or how the state of the system changed over time.

All of that context matters because incidents rarely happen in isolation. One event leads to another, which triggers something else, and eventually the problem manifests itself as the symptom that gets everyone’s attention.

Piecing together that timeline to find the root cause takes time, patience and, more importantly, experience. If some of that data is missing, figuring out what actually happened becomes harder.

It is the same with an SRE Agent.

We can build intelligence into the agent based on our own experience operating production systems, but the agent still needs access to the relevant data points to reconstruct the incident timeline and reason about the root cause.

Highly distilled, high-signal data is useful for detecting that something is wrong. But it doesn’t always contain enough context to explain why.

As we built our autonomous SRE Agent, one thing became increasingly clear: more data resulted in better outcomes than less data.

But the economics of observability push us in the other direction

The high cost of observability has become a significant pain point for many organizations.

As telemetry volumes grow, so do the costs of ingesting, storing and querying that data. Organizations have responded by introducing sampling, filtering logs, limiting high-cardinality data and reducing retention periods to keep observability costs manageable.

You are effectively forced to decide which telemetry is worth sending to your observability platform, how much of it to send and how long to keep it.

This creates an interesting conflict as we move toward AI-driven observability.

AI wants context. Traditional observability economics encourage us to remove context.

AI changes the value of telemetry

An SRE Agent investigating a latency incident doesn’t necessarily need every log permanently stored in a centralized SaaS platform.

What it needs is access to the relevant information when it is reasoning about the incident.

Imagine an SLO has been breached because latency has suddenly increased. Before getting anywhere close to a root cause, an agent might go through a series of steps:

High latency detected → identify affected service → examine traces → inspect logs → check recent deployments → inspect Kubernetes events → compare resource utilization → form a hypothesis → gather evidence

One step informs the next.

A trace might point toward a downstream service. That might lead the agent to its logs. The logs might suggest something changed recently. That could lead to deployment history and Kubernetes events. Resource metrics might then provide another piece of evidence supporting or disproving the hypothesis.

The value comes from the agent being able to traverse the data as it investigates, not necessarily from having centrally ingested all of that data beforehand.

That led us to an important shift in how we think about telemetry:

The question shifts from “What telemetry should we ingest?” to “What context should the agent be able to access?”

More context creates another problem

Of course, giving AI access to more operational data introduces another set of questions.

Enterprises need to understand where that data goes. What gets sent to an LLM? Does telemetry cross regions? Could sensitive information leave the environment? What permissions does the agent have? Can it retrieve information it shouldn’t? What information is retained? And what happens when an agent starts combining multiple tools as part of an investigation?

These aren’t just AI governance questions. They become data architecture questions.

Autonomous observability wants more context at exactly the time enterprises are becoming more concerned about data sovereignty and security.

Simply sending even more telemetry to a centralized platform so an AI agent can reason over it doesn’t really solve the problem. In some ways, it makes the problem bigger.

Move the intelligence to the data

This is where edge processing becomes interesting.

Instead of assuming that all telemetry needs to move to a centralized platform before intelligence can be applied to it:

Production → Telemetry → Centralized Platform → AI

we can move toward an architecture where processing and intelligence operate much closer to the source:

Production → Local Telemetry/Data Plane → Local Processing/Agent → Findings and Context → Control Plane

The idea is relatively simple:

Move intelligence closer to the data rather than moving all the data closer to the intelligence.

This gives organizations much more control over the boundary between their operational data and the AI systems consuming it.

They can decide what data an agent is allowed to access, what information can leave the environment, which models or LLM providers can receive that information, and what needs to remain local. Where the requirements justify it, models can also be run within the organization’s own environment.

This allows organizations to define data sovereignty and security on their own terms rather than having those decisions dictated by the architecture of their observability platform.

There is an economic benefit as well.

If the raw telemetry can remain close to where it is generated, organizations don’t necessarily have to continuously send every metric, log and trace to a centralized observability vendor just in case it becomes useful later. That can reduce both observability ingestion costs and the infrastructure costs associated with moving large volumes of telemetry.

Edge processing doesn’t mean everything stays local

None of this means that every piece of information needs to remain at the edge.

There are things that naturally make sense to keep local: raw telemetry, detailed logs and traces, high-cardinality metrics, Kubernetes state, investigation execution and potentially sensitive operational context.

There are also things that can make sense centrally: incidents, findings, summaries, fleet-wide views, policies, configuration and insights that need to span multiple environments.

The important difference is that moving the data becomes an explicit architectural decision rather than the default.

The question becomes: What needs to move, rather than assuming everything needs to move.

From telemetry pipelines to reasoning pipelines

We need to start thinking beyond telemetry pipelines and toward reasoning pipelines.

Traditional observability has largely been built around a familiar flow:

Collect → Transport → Store → Query → Human investigates

Autonomous observability changes that model:

Detect → Gather context → Correlate → Reason → Produce evidence → Act

That is a fundamentally different way of thinking about how we use telemetry.

An SRE Agent is tireless. It can continuously consume and correlate large amounts of operational data, reconstruct what happened and reason across multiple signals. The goal isn’t to replace the human SRE. It is to give the SRE team an advantage by doing much of the time-consuming investigation before a human needs to get involved.

Instead of starting an incident with a dashboard and a blank query window, an engineer can start with a hypothesis, the supporting evidence and a timeline of what happened.

In some cases, the agent can go one step further and take pre-approved mitigating actions, significantly reducing Mean Time to Resolution.

This model creates a strong incentive to move processing closer to the edge. If the agent can access more of the relevant operational context, and access it quickly, it has more information available to reason about what actually happened.

The alternative is to ingest everything into a centralized vendor platform and then have the agent retrieve the raw telemetry back through APIs as it investigates.

That starts to look increasingly inefficient.

You pay to move and ingest the telemetry. You depend on APIs to retrieve the context the agent needs. You introduce another layer between the agent and the systems it is investigating. And you give up some control over where that operational data is stored and processed.

From an economics, speed, control and data sovereignty perspective, that architecture can put autonomous systems at a disadvantage.

The rise of AI doesn’t necessarily mean we need to centralize more telemetry.

It may actually create a stronger reason to process more of it at the edge.

Reasoning should move closer to the data

One of the most important things we learned while building our SRE Agent was that giving it more context generally resulted in better outcomes.

But that doesn’t mean the answer is to centrally ingest even more telemetry.

The challenge for autonomous observability is giving AI access to the context it needs without requiring organizations to give up control of their data.

Our thesis is that the architecture needs to change with the interaction model.

If AI is going to continuously investigate systems, correlate telemetry, reconstruct incident timelines and eventually take action, then it makes sense for much of that intelligence to operate close to where the data is generated.

The raw telemetry can remain local. The agent can investigate it where it lives. Findings, evidence, incidents and summaries can move where they need to move.

This changes the role of the centralized observability platform. It doesn’t need to be the place where every piece of telemetry is sent before anything useful can happen. It can become the control plane for coordinating intelligence that operates much closer to the systems being observed.

The goal isn't to centralize all the data so AI can reason over it. The goal is to bring the reasoning closer to the data.

That is an important distinction.

In the age of autonomous observability, data sovereignty isn’t just about where data is stored.

It’s about where intelligence runs, what it can access, what actions it is allowed to take, and what information is allowed to leave the environment.

This is the architecture underpinning Randoli’s approach to building autonomous SRE Agents.

And as AI becomes a bigger part of how we operate production systems, I think this distinction will become increasingly important.

See how Randoli applies this in practice.