Introduction

Every AI roadmap for IT operations runs into the same wall, the AI is only as good as the operational data (retrievable context) beneath it. Legacy monitoring systems already struggle with hybrid cloud, distributed applications and Kubernetes; they produce silos, excessive costs and rigid tooling. Now those same weaknesses block the next step, because agentic AI cannot reason over telemetry it cannot reach, trust or correlate.

A composable observability architecture solves both problems at once. By integrating telemetry pipelines with a unified visualization layer, organizations improve operational efficiency, reduce tool lock-in and unlock greater value from their existing and future telemetry data. And because the architecture treats telemetry as an owned, governed asset rather than a vendor byproduct, it doubles as the data foundation that AI-driven operations requires. Since we first published this architecture in early 2025, that second point has moved from a footnote to a headline.

Challenges with legacy monitoring

Common challenges with legacy monitoring solutions include:

  • Monitoring silos: Different teams use different tools, preventing seamless access and correlation of data.
  • Increasing costs: Traditional logging solutions keep logs within the vendor tools, often resulting in high storage costs with limited flexibility. Scaling to an enterprise comes with hidden arbitrage.
  • Burdensome migrations: Shifting to new monitoring tools is difficult, often requiring major overhauls to both process workflows and training.
  • Difficulties isolating outages: Legacy systems lack correlation between disparate telemetry sources, making root cause analysis a time-consuming effort.
  • Tool lock-in: Monitoring solutions impose rigid UIs, fixed storage formats and restricted data access, limiting flexibility and making it challenging to adopt new tools or modernize your observability stack.

To this list, the last 18 months have added a sixth challenge: AI readiness. Organizations want AI-assisted root cause analysis, conversational access to operational data and, increasingly, autonomous agents that act on telemetry. Every one of those capabilities inherits the flaws of the data layer it sits on. Fragmented, silod telemetry produces fragmented, unreliable AI.

The interconnection of assets, automation and observability

Let's take a quick step back and understand the IT landscape.  

A triangle with text on it

AI-generated content may be incorrect.

Assets, automation and observability are deeply interconnected, forming the foundation of modern IT operations.

Assets represent the infrastructure, applications, services and devices that generate telemetry data. Without a comprehensive understanding of these assets, organizations cannot effectively monitor or manage them. How do you monitor something you do not know you own? How can you monitor for something you know should not be present?

Automation relies on that same asset understanding so it knows what it can act upon. Automation is also essential for ensuring telemetry data is collected, processed and acted upon in real time, reducing manual intervention in remediation. The pull-through effect of automation is standard definitions expressed as code, which ensures uniform application of changes and repeatability.

Observability relies on both: without a well-managed asset inventory, observability lacks awareness of what to monitor, and without automation there is no mechanism to provision, change and remove what observability reveals.

An effective observability strategy requires an integrated approach where these three components continuously feed into and enhance one another. Keep this triangle in mind; it returns later in the article, because these same three components (a source of truth, deterministic automation and the observability to measure the system) are precisely the prerequisites for agentic operations.

Understanding the composable observability architecture

A composable observability architecture consists of five layers:

  1. Telemetry producers. The devices and systems that generate raw telemetry data: metrics, logs, events and traces. These are the sources of data, such as servers, routers, applications, IoT devices and services.
  2. Telemetry pipeline (observability pipeline). The backbone for routing, enriching and transporting telemetry data efficiently to various destinations. It decouples data generation from consumption, allowing multiple tools to ingest and analyze data. The pipeline is designed to preserve data integrity, apply necessary transformations and enrichment, and provide a scalable mechanism for telemetry distribution across the organization. The pipeline also functions as a service bus that enables event-driven architecture: decoupling not only producers from consumers but also the telemetry consumers from one another.
  3. Telemetry consumers. The analytics, security and operational platforms that derive actionable insights from telemetry data: network monitoring systems (NMS), application performance management (APM) platforms, logging platforms and others. Consumers depend on a structured, well-managed pipeline for accurate and timely access to telemetry.
  4. Analytics layer. Responsible for deriving insights from telemetry data and from the analytics conducted by downstream tools, applying event correlation, policy management, event analysis and AI/ML-driven functions. It enhances event management by prioritizing alerts, identifying patterns and predicting failures before they impact operations. This layer enables root cause analysis, anomaly detection and automated response actions.
  5. Visualization layer. The interface that provides end users with a clear, unified view of observability data. By consolidating data from multiple telemetry consumers, it ensures different IT teams (network engineers, service desk analysts, application owners) work from a shared, reliable perspective. It minimizes operational burden with persona-based dashboards and enables tools to be added or decommissioned without disruptive changes to workflows, interfaces, training or SOP documents.
Layered observability architecture showing telemetry collection, data management, analytics, and enterprise services
Composable Observability Architecture 

How AI accelerates every layer

A fair question in 2026: where does AI fit in this architecture? The answer is not a sixth layer or a new box on the diagram. AI is an accelerator inside every layer that already exists. That is the point of composability: the architecture absorbs new capabilities without being redrawn.

  • At the producers, AI-assisted instrumentation is reducing the manual work of getting telemetry out of systems in the first place: recommending collection configurations, identifying under-instrumented assets against the source-of-truth inventory and flagging telemetry gaps before they become blind spots during an incident.
  • In the pipeline, AI augments transformation and enrichment. Machine learning can classify and route unfamiliar log formats, recommend aggregation and sampling rules that cut storage cost without losing signal, and detect drift in data quality (a producer that silently changes its log schema, for example) before downstream consumers break.
  • At the consumers, most major observability platforms now embed AI-driven analysis natively, from automated baselining to anomaly detection to assisted investigation. Because the pipeline feeds every consumer from the same governed stream, these AI features work from consistent data rather than each tool's partial view.
  • In the analytics layer, AI does its most visible work: correlating events across domains, compressing alert floods into a small number of probable-cause incidents, and assembling root cause narratives as incidents unfold rather than after the fact.
  • In the visualization layer, conversational interfaces are changing who can use observability data. A service desk analyst can ask, in plain language, "what changed on the payment service in the last hour?" and get an answer synthesized from metrics, logs, traces and change records, without knowing which underlying tool holds each piece.

Notice what every one of these accelerators has in common: each consumes broad, consistent, well-governed telemetry. AI embedded in a siloed tool can only be as insightful as that tool's slice of the data. AI operating on a composable architecture sees the whole system. The architecture does not just tolerate AI; it is what makes AI worth deploying.

The value of the telemetry pipeline

The telemetry pipeline offers significant advantages:

  • Data access and reuse. A centralized telemetry pipeline eliminates duplicated data collection, making telemetry data, both real-time and historical, accessible to multiple teams across the organization. All departments, from IT operations to security to development, work from the same reliable data source.
  • Event-driven architecture. A decoupled system design allows applications to operate independently while communicating through event streams. This differs from traditional request-response architectures, where service dependencies create bottlenecks and single points of failure. Operations teams benefit from enhanced scalability, recoverability from system failures without data loss, and asynchronous processing.
  • Centralized telemetry management. By serving as the authoritative distribution point for telemetry data, the pipeline promotes consistency and eliminates data fragmentation and silos of knowledge across teams.
  • Flexible data storage. Unlike legacy monitoring solutions that impose rigid storage structures, a telemetry pipeline allows a flexible storage approach: on-premises, cloud or hybrid, optimized for cost and accessibility. Telemetry data can be transformed to conform to the formats each consuming application expects.
  • Loosely coupled integrations. Traditional monitoring solutions often require tight integration between producers and consumers, making it difficult to adopt new tools. A telemetry pipeline introduces loose coupling, allowing IT teams to switch or add observability tools without major infrastructure changes or risky cutovers.
  • Historical state and recovery. A well-designed pipeline retains historical telemetry across all producers, allowing operations teams to conduct forensic analysis, track trends and investigate long-term performance degradations.
  • Vendor independence. Many OEM vendors lock organizations into proprietary ecosystems and data structures, limiting flexibility and driving up costs. A telemetry pipeline keeps data ownership within organizational control; telemetry is not held captive inside a proprietary vendor tool or an external partner's system.
  • An AI-ready corpus. This is the advantage that has grown most since we first published this architecture. Every AI initiative in IT operations, from assisted root cause analysis to autonomous agents, needs exactly what the pipeline provides: broad, consistent, historical, well-governed telemetry in formats the organization controls. Teams that built pipelines for cost and flexibility reasons discovered they had also built their AI data foundation. Teams that skipped this step are now retrofitting it before their AI investments can deliver.

Telemetry pipeline example

To illustrate how a composable observability architecture functions, here is a real-world example. The specific tools reflect one deployment; the pattern (collect, stream, consume, correlate, visualize, automate) is what matters, and it has held up as individual products have added AI capabilities around it.

End-to-end observability and event streaming architecture connecting telemetry sources, analytics platforms, and automation tools
Example:  Real-World Observability Architecture 

Telemetry producers consist of network devices, applications and infrastructure components generating telemetry data, including SNMP metrics, logs, traces and events.

Telemetry pipeline collects, processes and transports telemetry data to downstream consumers.

  • Grafana Alloy
    • Polls network devices via SNMP to gather network performance metrics.
    • Uses remote_write to send structured time-series data to Cribl Stream over HTTP for further processing.
  • Cribl Stream
    • Ingests logs, metrics and traces from various sources, including syslog via HTTP, TCP and UDP.
    • Optimizes log data through aggregation and deduplication, reducing storage costs while maintaining valuable insights.
    • Routes processed telemetry to Confluent Cloud Kafka, categorized into specific telemetry topics.
  • Confluent Cloud
    • Acts as the central service bus for telemetry data, enforcing a "write once, read many" model.
    • Exposes Kafka topics for different telemetry types (metrics, logs, traces, events) to provide structured access for observability tools.
    • Uses processing rules to enrich and transform data before exposing it to consumers.
    • Supports event-driven integration among analytics, correlation, incident, alerting and notification systems.

Telemetry consumers and point tools. Tools interact with the pipeline in two directions. Some genuinely consume raw telemetry from the pipeline; others collect their own telemetry directly from the environment and instead publish their distilled output (events, alerts, analyzed problems) into the pipeline for downstream correlation and automation. In both cases, the pipeline is the integration point; no tool needs a point-to-point connection to any other.

  • NMS, NPM and security tools
    • Low-level collectors feed the pipeline with raw network telemetry; higher-level network management and performance monitoring point tools ingest from the environment through their own collection mechanisms and analyze it for network health, traffic optimization and SLA monitoring.
    • Security tools, including SIEM platforms, consume logs and events from Kafka topics for event-driven security analytics and real-time threat detection.
    • All of these tools publish their processed events and alerts into Confluent Cloud, where correlation, analysis and automation tools consume them.
  • Dynatrace (APM)
    • Collects application telemetry directly through its own agents and gateways for deep observability into applications, microservices and distributed environments. It can also ingest OpenTelemetry-format data from other sources through this collection tier.
    • Applies AI-powered analytics to detect performance bottlenecks and anomalous behavior proactively.
    • Publishes the results of that analysis (problems, alerts, events) into Confluent Cloud for correlation in BigPanda and automated response workflows. Raw application metrics, logs and traces stay in the APM platform; routing them through the pipeline would add cost without adding value.

Analytics layer provides event correlation, deduplication and automated incident generation, ensuring only meaningful alerts escalate to ITSM.

  • BigPanda
    • Acts as the event correlation and analytics platform authorized to generate incident tickets in the ITSM tool.
    • Consumes alerts from multiple observability tools via Confluent Cloud and applies AI-driven event correlation and analysis.
    • Compresses, deduplicates and prioritizes alerts into actionable incidents, significantly reducing noise and false positives.
    • Sends correlated events back into Kafka topics for consumption by the ITSM system and to initiate automated remediation via Ansible Automation Platform.

Visualization layer, powered by Grafana Cloud, provides unified dashboarding and reporting across all observability components.

  • Grafana Cloud
    • Serves as the enterprise-wide reporting and dashboarding solution, visualizing key metrics, logs, traces and event analytics from multiple sources.
    • Provides persona-driven dashboards tailored for network engineers, security analysts, application owners and IT leadership.
    • Enables trend analysis, reporting and SLA tracking across all solution components.

Automation architecture. Although not depicted in the diagram, the entire solution is delivered through infrastructure as code (IaC), powered by Red Hat OpenShift and Ansible Automation Platform.

  • Red Hat OpenShift
    • The telemetry collectors (Cribl and Grafana Alloy) have their configurations and manifests managed in Git for source control, then deployed onto distributed OpenShift clusters using GitOps methodology with ArgoCD.
  • Ansible Automation Platform
    • As much of the solution as possible is managed as IaC. Confluent integrations, Kafka topics and Grafana dashboards are managed through Git source control and CI/CD pipelines and deployed by Ansible Automation Platform. This enables rapid provisioning or reconfiguration of any component.
    • Listens to Kafka topics on Confluent, triggering automated remediation actions based on observed telemetry: infrastructure adjustments, service restarts or configuration changes. Ansible can also create or update incident tickets within ITSM platforms, ensuring tracking and process compliance.

From observability to agentic operations

The industry conversation has moved past AI that recommends toward AI that acts: agentic operations, where autonomous agents detect a signal, investigate across telemetry, form a root cause hypothesis and execute or propose remediation, escalating to a human when judgment is required. Two use cases are leading the way in practice: AI-assisted root cause analysis, where agents assemble probable-cause narratives from correlated telemetry as incidents unfold, and conversational operations, where chat interfaces give engineers and service desk analysts natural language access to operational data and runbook actions. WWT has explored both across our observability and AIOps work and in agentic AIOps testing in our Advanced Technology Center.

Here is the uncomfortable truth behind the excitement: an agent is only as trustworthy as the system it operates within. Before an autonomous agent can safely act on your environment, three things must already be true. They should look familiar; they are the same triangle this article opened with.

  1. A source of truth. The agent must know what exists, what is intended state and what "normal" looks like. That requires the well-managed asset inventory and governed telemetry that the composable architecture provides. An agent reasoning over incomplete or contradictory data will act confidently and wrongly, which is worse than not acting at all.
  2. Deterministic automation. Agents should not improvise change. They should select from tested, version-controlled automation (the same IaC, GitOps and Ansible patterns described above) so that every action an agent takes is one the organization has already defined, reviewed and made repeatable. The AI decides when and which; the deterministic layer defines what and how.
  3. Observability to measure the system, including the agents. Every agent action is itself telemetry. It flows through the same pipeline, is correlated in the same analytics layer and appears on the same dashboards as any other change to the environment. If you cannot observe your agents, you cannot trust them, tune them or turn them off with confidence.

Organizations that have implemented a composable observability architecture will recognize this list: it is the architecture. That is the strategic payoff of the approach. The work done to decouple telemetry, codify automation and unify visibility was never just a monitoring modernization; it built the operating foundation that agentic AI requires. Organizations skipping ahead to agents without that foundation are automating on sand.

Guardrails: policy as code

Autonomy without constraint is a liability, and the guardrail question "what stops the agent from doing something destructive?" deserves a concrete answer, not reassurance.

Our answer is policy as code, enforced through deterministic automation against the source of truth. Policies (what actions are permitted, on which assets, under which conditions, with what blast radius and approval requirements) are defined in version-controlled code, just like the infrastructure itself. Agents do not receive open-ended access to the environment; they receive access to a catalog of pre-approved, tested automation actions, each governed by policy. When an agent determines a service restart will resolve an incident, the restart it triggers is the same Ansible playbook a human operator would run, executed under the same policy checks, logged into the same audit trail and visible in the same dashboards.

This model has three properties worth calling out:

  • Auditability. Every agent decision and action is recorded as telemetry, reviewable after the fact and correlated with its outcome.
  • Graduated autonomy. Policy defines which actions run autonomously, which require human approval, and which are recommendation-only. Organizations can start conservative and expand agent authority as measured trust grows, per action class, not as a single on/off switch.
  • Reversibility. Because actions are drawn from version-controlled automation against a known source of truth, the system can reason about rollback, and humans can see precisely what changed.

Guardrails, in other words, are not a bolt-on safety feature. They are the same disciplines (as-code definitions, source of truth, observability) applied to a new class of operator.

The Unified Visualization Layer

Imagine an IT operations director has just been told they must move to a new network monitoring tool, requiring an overhaul of how the service desk operates. The new tool has a different UI, different alerting mechanisms and an entirely new way of viewing network health. The team has spent years developing efficient workflows, writing standard operating procedures (SOPs) and ensuring everyone, from the day shift to the night shift, knows how to triage issues. Now all of that is about to change.

This is where a unified visualization layer transforms the experience. Instead of forcing the service desk to abandon existing processes and dashboards, the visualization layer integrates the new system's data into the familiar dashboards the team already uses. The new tool plugs into the pipeline, the old tool is removed, and the service desk's dashboards, processes and workflows carry forward with little or no disruption. The service desk works with a tailored visualization layer that abstracts the plumbing underneath.

A screenshot of a computerAI-generated content may be incorrect.
Example:  Grafana Labs Dashboard by Personas 

The power of persona-based dashboards

  • Role-specific views. A service desk analyst sees different data than a network engineer or an application owner, reducing irrelevant alerts and improving efficiency.
  • Noise reduction. By filtering out unnecessary information, analysts focus on what's important, minimizing alert fatigue.
  • Operational consistency. Even as new monitoring tools are introduced, the visualization layer ensures IT teams continue to interact with data in a familiar, structured way.
  • Faster decision-making. Dashboards highlight key performance indicators, enabling engineers to diagnose and resolve issues more efficiently.

IT operations teams can integrate multiple tools seamlessly, avoid massive workflow disruptions and ensure operational continuity. This approach saves time, resources and frustration, allowing IT teams to focus on proactive network management rather than navigating disruptive tool changes.

The dashboard is also no longer the only front door. Conversational interfaces backed by AI are becoming a second access pattern for the same data: an analyst asks a question in natural language and receives an answer drawn from across the telemetry estate. This works for the same reason the persona dashboard works. The visualization layer, dashboard or chatbot, is only as good as the breadth and consistency of the data behind it, and that is what the pipeline provides.

Conclusion

The composable observability architecture, powered by telemetry pipelines and a unified visualization layer, changes how IT operations teams manage complexity, efficiency and resilience. By breaking down data silos, reducing vendor lock-in and enabling seamless integration of monitoring tools, organizations build a flexible, scalable observability framework that evolves with their needs.

What has changed since we first published this architecture is the stakes. Owning your telemetry, codifying your automation and unifying your operational view used to be an efficiency argument. It is now the entry requirement for agentic operations. The organizations that made these investments are positioned to adopt AI-driven operations deliberately and safely; the rest will find that no AI platform, however capable, can compensate for a fractured data foundation.

Start where the leverage is: take control of your telemetry. The rest of the architecture, and everything AI will build on top of it, follows from there.

Technologies