Topics in this article

Telecommunications networks are highly distributed and dynamic environments where service quality can begin to deteriorate long before conventional alarms are raised. As traffic volumes grow and architectures become more distributed, traditional monitoring approaches are struggling to keep pace, creating a need for more adaptive and intelligent approaches.

At NTT DATA, we’re developing agentic AI solutions powered by NVIDIA technologies to transform network operations from reactive monitoring into continuously adaptive and autonomous decision-making systems.

A key enabler is NVIDIA® NemoClaw™, a blueprint or reference stack for deploying autonomous agents (such as OpenClaw) with open harnesses, model access and OpenShell controls.

Combined with NVIDIA NIM™ microservices, agentic workflow frameworks and next-generation operations support system (OSS) alarm-management workflows, NemoClaw can reason across large volumes of time-series telemetry and intelligently coordinate tools and actions throughout network operations.

From reactive monitoring to continuous reasoning

Traditional OSS monitoring relies on static thresholds and reactive alarms, which are increasingly insufficient to capture complex, evolving network behaviors in distributed environments.

NemoClaw, together with NVIDIA Nemotron™ models, brings real-time intelligence to streaming telemetry. By analyzing network behavior in real time, AI agents can detect weak signals and track evolving anomalies across long-term key performance indicator (KPI) patterns.

Intent-based routing enables dynamic switching between shallow and deep reasoning paths, while secure execution and governance controls can be enforced through NVIDIA OpenShell runtime, providing sandboxed execution, policy enforcement, network guardrails and privacy-aware data routing for agentic workflows.

This approach can also help manage inference cost by reserving deeper reasoning paths only for relevant anomalies. As a result, it reduces unnecessary use of large language models and supports scalable agentic operations.

A use case: Preventing silent network degradation

Silent network degradation refers to progressive performance issues — such as gradual latency increases, throughput drops or localized quality deterioration — that traditional threshold-based monitoring systems often don’t detect until the user experience is affected.

These issues typically emerge over time in distributed network domains, making them difficult to identify through isolated alarms or static rules.

To address this challenge, we have implemented an agentic architecture powered by NVIDIA NemoClaw, combining adaptive orchestration, intent-based routing and multilayer telemetry analysis.

Instead of relying on isolated alerts, the system analyzes network dynamics holistically and uses hierarchical reasoning and multiagent workflows to investigate suspicious patterns. Through structured reasoning, it detects and contextualizes silent degradations and gives first-line operations teams clear visibility of affected services, probable root causes and recommended remediation actions.

As a result, operators can help reduce manual investigation effort, accelerate response times and shift from reactive troubleshooting to proactive network operations.

A diagram showing interaction between NTT DATA NextGen OSS and NVIDIA NemoClaw

End-to-end flow for proactive silent network detection

The architecture combines multiple telemetry layers. Real-time telemetry streams provide fine-grained operational visibility, while aggregated KPI analytics enable scalable long-term trend analysis and historical reasoning across the network.

How it works

  • At the core, an anomaly agent orchestrated through NVIDIA NemoClaw uses a lightweight NVIDIA Nemotron Nano model (Nemotron Nano 12B v2 VL) served through NVIDIA NIM microservices to continuously monitor aggregated KPIs across the network, focusing on long-term behavioral trends (such as one-hour windows) to detect early degradation signals, including latency increases, throughput decline or localized quality drops. This layer provides scalable observation across the network. When anomalies are detected, the system escalates only relevant cases into deeper analysis workflows. This hierarchical reasoning approach reduces computational overhead.
  • A research agent orchestrated through NVIDIA NemoClaw and powered by an NVIDIA Nemotron Super model (Nemotron-3-Super-120B-A12B) performs high-resolution investigation using fine-grained telemetry (15-minute or sub-interval granularity). The model is optimized for complex analytical tasks focusing on affected regions to provide contextual comparison and determine whether the issue is isolated or relates to broader network behavior.
  • Then, an enrichment agent uses tool orchestration and retrieval workflows to enrich anomalies with operational context from OSS systems — including topology, inventory and service dependencies — and translate telemetry anomalies into operationally meaningful events.
  • Once enriched with operational context, each degradation is evaluated by the prioritization agent, which weighs factors such as anomaly severity, duration, equipment context and related alarms to assign an impact score. First-line operations teams can direct their attention where it matters most, which reduces response times on high-impact incidents.
  • Finally, the recommendation agent consolidates prioritized findings, evaluates probable root causes and suggests remediation actions, supporting faster operational responses.

The workflow ensures focused investigation by combining shallow monitoring with targeted deep analysis while maintaining contextual awareness across the network.

How telco network operators can benefit

When operators combine NemoClaw, intent routing and next-generation OSS integration, they gain a foundation for autonomous, always-on network operations. The benefits include:

  • Faster identification and isolation of network degradation patterns
  • Reduced manual effort in triage and investigation workflows
  • Improved decision quality through contextual and multilayer analysis
  • Better operational visibility across distributed network domains
  • Integration with existing OSS and monitoring ecosystems
  • A clear path toward closed-loop autonomous operations through the integration of automated corrective actions and policy-driven remediation workflow
  • Strong governance and operational safeguards for AI agents, which operate within OpenShell-style policy controls and operator-defined trust boundaries. A privacy router keeps subscriber and operational telemetry within the operator’s trusted environment.

Redefining telco network operations

Our expertise in telco operations and agentic solutions, combined with NVIDIA’s agentic AI stack, helps network operators move from reactive network management to systems capable of continuous reasoning.

With AI agents supporting operators with telemetry analysis and recommended actions, networks can now observe, interpret and investigate themselves in real time while keeping operational costs under control.

This approach lays the foundation for a new generation of autonomous, scalable and economically sustainable networks.

WHAT TO DO NEXT
Learn more about our success with NVIDIA and contact us to explore how we can partner with you on your AI transformation journey.