You're reading the Saturday, October 3, 2026 edition. Today's briefing →
Toronto
Saturday, October 3, 2026No. 111

Digital Plumber

Plumbing the information age

AI-curated intelligence for people who run networks. Daily coverage of AIOps, network automation, agentic operations, AI infrastructure, security and the vendors shaping them.

Today's 3 things that matter

Picked by the AI editor
  1. Telco·Analysis

    Verizon moves from scripted automation to agentic network operations

    Verizon is moving from scripted RAN automation to agentic network operations where AI reasons across domains and takes action while operators retain control of intent and guardrails.

    Why it matters Verizon distinguishes vendor domain agents from operator-owned orchestration layer; operators must own multi-domain intelligence because only they hold full topology knowledge, preventing uncontrolled autonomous behavior that could cascade across RAN, transport and backhaul.

  2. Routing·Research

    RPKI is not a significant source of BGP noise

    APNIC analysed 11 years of RPKI and BGP data to measure routing activity associated with RPKI-related changes.

    Why it matters Operators deploying Route Origin Validation can proceed with confidence that RPKI deployment will not destabilize the routing system at current adoption rates.

  3. Research·Primary source

    IETF draft proposes governance for AI-run network devices

    Internet-Draft draft-smith-opsawg-ai-network-governance-01 proposes a governance framework for AI systems (particularly LLMs) autonomously managing network device operations.

    Why it matters NetOps teams adopting AI-driven automation must consider governance, risk, and compliance models; this draft provides IETF-level guidance for safe LLM deployment in autonomous network management workflows.

Today's briefing

What happened, and why it matters

9 stories · 7 topics · Updated 7:19 PM ET

LogicMonitor claims 80% alert noise cut with agentic AIOps

LogicMonitor platform page · Sep 30, 2026 · Primary source

LogicMonitor's LM Envision platform helps IT teams eliminate alert storms, enrich events with AI context, and automate issue resolution, claiming 65% faster incident resolution and 80% reduction in alert noise. Platform unifies network, infrastructure, and cloud observability with embedded agentic AIOps.

Why it matters Cloud monitoring delivers more value when environments share the same observability context, giving teams clearer troubleshooting, cost management, and support for AI-assisted operations—critical for hybrid NOC consolidation.

LogicMonitor's LM Envision represents convergence: unified observability across network, infrastructure, and cloud with agentic AIOps embedded in the correlation layer. The 65% MTTR improvement and 80% alert noise reduction reflect maturation of event correlation beyond static rules—the platform correlates alerts, metrics, logs, and topology to suppress downstream symptoms and surface root causes. Notably, LogicMonitor supports synthetic monitoring and Internet Performance Monitoring (IPM), extending observability beyond the enterprise boundary to detect cloud and SaaS failures before customers do. The platform positions itself as addressing the NOC's growing scope: not just network and servers, but SaaS uptime, API performance, and user experience. For operations teams managing multi-tenant or managed services, agentic capabilities include automated remediation gates, dynamic thresholds with seasonality and rate-of-change detection, and root cause analysis spanning network, infrastructure, and cloud layers—reducing manual escalation runbook overhead.

Read the original at logicmonitor.com ↗

HPE lifts AI networking order target to $3 billion

Yahoo Finance · Oct 2, 2026 · Industry news

HPE raises fiscal 2026 Networks for AI order target to $3 billion. Post-Juniper integration shows 117% networking revenue growth YoY through July 2026; networking now 26% of HPE segment revenue. QFX5250 switches on Broadcom Tomahawk 6 (102.4 Tbps) and MX301 routers shipping to cloud and hyperscaler customers. HPE targeting $800M cost synergies by FY28.

Why it matters Demonstrates integrated routing/switching/telemetry stacks command premium pricing in AI infrastructure. Distributed AI workloads and data center interconnect drive routing growth at 20%+ CAGR, shifting network capex composition toward layer 3 intelligence and multi-vendor scale-out fabrics.

HPE's October 2026 earnings guidance update reflects the strategic impact of the Juniper acquisition (closed July 2025). Networking revenue grew from 15% of segment sales (pre-Juniper) to 26% of $32 billion total segment sales in the nine-month period ending July 31, 2026. HPE expects routing revenue to grow at 20-30% CAGR through fiscal 2029, driven by AI on-ramp applications connecting users/branches to AI infrastructure and multi-region data center interconnect (DCI). The company is participating in Oracle's multiyear gigawatt-scale AI infrastructure buildout with MX and PTX routers supporting edge and regional fabric layers, while QFX switches power multiple layers of regional data center fabrics. Marvis self-driving AI (from Juniper Mist platform) is extending into HPE Aruba Central for unified campus and data center network operations. QFX5250 switches support both Broadcom Tomahawk 6 silicon and open-standard UALink interconnect as an alternative to NVIDIA proprietary NVLink, offering customers architectural flexibility for hybrid compute clusters. Announced cost-synergy target of $800M by FY28 (raised from $600M at deal close) supports high-teens revenue CAGR and mid-to-high 20s operating margins through FY29. This reflects conviction that AI infrastructure networking will remain a sustained high-growth category separate from traditional enterprise switching.

Read the original at finance.yahoo.com ↗

RPKI is not a significant source of BGP noise

APNIC Blog · Oct 2, 2026 · Research

APNIC analysed 11 years of RPKI and BGP data to measure routing activity associated with RPKI-related changes. Study finds RPKI is not currently a significant source of BGP noise despite growing ROA adoption and ROA management activity.

Why it matters Operators deploying Route Origin Validation can proceed with confidence that RPKI deployment will not destabilize the routing system at current adoption rates.

Each ROA creation, modification, revocation, or expiration triggers routers performing ROV to reconsider routing decisions, potentially propagating through BGP and generating additional updates. APNIC reconstructed RPKI events from available datasets and identified BGP updates during convergence periods to quantify RPKI-induced routing activity. The number of ROA creations, expirations, and revocations has increased significantly over time for both IPv4 and IPv6, reflecting steady RPKI adoption. However, the research demonstrates that despite growing ROA management activity, RPKI changes do not currently contribute materially to global BGP churn. The study used an 11-year time machine methodology comparing RIPE Archive and Flutter data to establish reliable historical estimates. Results show that RPKI introduction of new routing-change triggers has not exceeded operator capacity to process updates, and the routing system remains stable under current RPKI deployment levels.

Read the original at blog.apnic.net ↗

Race condition with schema migration broke Railway routing globally

IncidentHub / Railway · Sep 30, 2026 · Primary source

Railway's global routing service returned 404 errors for five minutes after a race condition caused a new routing service to deploy before its required database schema migration completed. Service queried non-existent columns, causing all route lookups to fail across all regions.

Why it matters Illustrates critical deployment ordering and validation challenges for distributed routing control planes; demonstrates why schema validation must precede service rollout in mission-critical infrastructure.

On September 30 at 07:29 UTC, an engineer merged a change to Railway's routing control plane including both a database schema update and new routing service version. A race condition in the pipeline configuration ran the schema migration in parallel with service deployment instead of serially. At ~07:35 UTC the new routing service rolled out across all regions and immediately began querying columns that did not yet exist, causing every route lookup to fail. The edge treated lookup failures as non-existent domains rather than falling back to cached routes, returning 404s for all customer domains including *.up.railway.app addresses. Internal alerting notified the on-call team; at ~07:40 UTC the CI pipeline completed the schema update, the routing service retried lookups successfully, and domains recovered. Total blast radius: all four Railway regions for five minutes affecting approximately 3 million users. Railway's remediation includes enforcing serial ordering between schema migrations and deployments, blocking service rollouts until schema validation passes, and implementing shard-based regional rollouts to limit future blast radius.

Read the original at blog.incidenthub.cloud ↗

Operators rebuild fiber and data centers for AI workloads

TelecomLead · Sep 30, 2026 · Industry news

Telecom operators are rebuilding networks for the AI era as GPU clusters, distributed inference and AI data centers create new requirements for fiber capacity and low-latency connectivity. AT&T's $3 billion September 2026 Corning agreement will supply fiber for both broadband expansion and long-haul AI data center interconnection.

Why it matters AT&T and Verizon securing massive fiber volumes, SK Telecom targeting 15 GW AI data-center capacity; fiber investment now dual-purpose infrastructure supporting consumer broadband and hyperscaler AI connectivity, reshaping capex priorities and competitive positioning.

Telecom operators are rebuilding networks for the AI era as GPU clusters, distributed inference, cloud applications and AI data centers create new requirements for fiber capacity, optical transport, computing and low-latency connectivity. The shift is changing telecom capex allocation. Operators are no longer investing only in spectrum, radio networks and consumer broadband. AT&T and Verizon are securing massive volumes of fiber; SK Telecom is targeting 15 GW of AI data-center capacity; Airtel is expanding toward 1 GW of data centers; and SoftBank is combining GPU clouds with AI-RAN. AT&T's September 2026 agreement with Corning, valued at more than $3 billion, will supply fiber and cable for network expansion serving two purposes: expanding broadband to homes and businesses and creating long-haul routes connecting AI data centers. The emerging network architecture increasingly connects AI data centers → long-haul fiber → metro networks → edge computing → 5G → devices. For infrastructure practitioners, the key implication is that fiber capex is becoming dual-purpose infrastructure supporting both consumer broadband and hyperscaler AI connectivity, fundamentally reshaping competitive positioning, M&A rationale and vendor relationships.

Read the original at telecomlead.com ↗

Verizon moves from scripted automation to agentic network operations

VoIP Review · Oct 2, 2026 · Analysis

Verizon is moving from scripted RAN automation to agentic network operations where AI reasons across domains and takes action while operators retain control of intent and guardrails. The carrier processed 70 million configuration changes in 2025 via closed-loop automation and now targets Level 4 autonomy across domains.

Why it matters Verizon distinguishes vendor domain agents from operator-owned orchestration layer; operators must own multi-domain intelligence because only they hold full topology knowledge, preventing uncontrolled autonomous behavior that could cascade across RAN, transport and backhaul.

Verizon is moving from scripted RAN automation towards agentic network operations, where AI reasons across domains, shares context and takes action with the operator retaining control of intent, outcomes and guardrails. Verizon wants to move beyond rule-based automation toward AI systems that can reason about problems and act across RAN, transport and other network domains. The carrier already uses closed-loop automation at large scale. In 2025, its platforms processed more than 70 million configuration changes, saving technician hours and reducing manual effort. This marks a shift from scripted network tasks to agentic operations. Verizon argues that vendors can supply powerful RAN agents, but only the operator has the network topology and institutional knowledge to arbitrate between them. The company is working towards more autonomous closed-loop action, but wants humans to define intent, control consequential changes and retain oversight of major incidents. A radio issue may originate from transport, backhaul or software behavior; a single vendor agent may miss that wider picture. For network operations leaders, this represents a critical architectural principle: vendor AI agents handle tactical domain problems, but operators must own the multi-domain orchestration and reasoning layer to prevent cascading failures across interdependent infrastructure.

Read the original at voip.review ↗

IETF draft proposes governance for AI-run network devices

IETF Datatracker · Sep 27, 2026 · Primary source

Internet-Draft draft-smith-opsawg-ai-network-governance-01 proposes a governance framework for AI systems (particularly LLMs) autonomously managing network device operations. The framework addresses detection, diagnosis, and remediation of operational anomalies on network infrastructure using AI services.

Why it matters NetOps teams adopting AI-driven automation must consider governance, risk, and compliance models; this draft provides IETF-level guidance for safe LLM deployment in autonomous network management workflows.

Published September 27, 2026, this individual Internet-Draft (draft-smith-opsawg-ai-network-governance-01) from Arrcus engineers addresses a critical gap in how autonomous network management using large language models should be governed. The draft defines a governance framework for systems that use AI services—specifically LLMs—to autonomously detect, diagnose, and remediate operational anomalies on network devices. Unlike prescriptive technical specifications, this informational draft takes a framework-based approach to governance, addressing the intersection of AI autonomy, network operations, and risk management. The framework is intended for IETF working group opsawg (Operations and Management Area Working Group), targeting practitioners building AIOps and autonomous remediation systems. This signals growing IETF focus on LLM safety and governance in network operations contexts. The draft expires March 31, 2027, allowing for community feedback and potential evolution into more formal guidance.

Read the original at datatracker.ietf.org ↗

Google's Gemini 4 Argon patches vulnerabilities autonomously

The Hacker News · Oct 1, 2026 · Industry news

Google unveiled Gemini 4 Argon, a frontier model with 1M-token output limit (up from 64K) for complex reasoning tasks in software engineering, enterprise knowledge work, and cybersecurity. The model can autonomously discover, validate, and patch critical vulnerabilities.

Why it matters Argon rolls out first to cybersecurity defenders through Fairwind Program with guardrail-free version capable of autonomous vulnerability discovery and patching—critical capability for SRE teams managing infrastructure security automation.

Google unveiled Gemini 4 Argon on October 1, 2026, a new frontier model designed for complex reasoning tasks requiring sustained thought. The model's key technical advancement is expansion of output capacity from 64,000 tokens to 1 million tokens, enabling agents to work through lengthy and complicated problems in a single run. This headroom lets the model think and write for hundreds of thousands of tokens continuously. Argon delivers frontier performance in real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. For cybersecurity specifically, the model can autonomously find, validate, and patch critical software vulnerabilities. Google is rolling out Argon first to trusted cyber defenders through its Fairwind Program, releasing a version without cyber guardrails to those defenders and internal teams. Broader access for paid API customers and Google AI Ultra subscribers is planned but not yet dated. Introductory pricing is $2 per million input tokens and $10 per million output tokens. Google reported that internally, Argon improved quantum computing efficiency by 40% and AI agent teams identified data-center memory optimizations freeing 300+ tebibytes of capacity. A notable finding: through Wiz's Scan for Good program, Argon uncovered a critical vulnerability in healthcare software used by hospitals worldwide that earlier frontier models had missed.

Read the original at thehackernews.com ↗

UK publishes AI risk toolkit with nine risk categories

TLT's AI Brief · Oct 3, 2026 · Primary source

The UK government published a new AI Risk Management Toolkit establishing nine categories of AI risk including legal compliance, fairness, transparency, accountability, and technical robustness. The toolkit explicitly cross-references existing legal obligations like data protection and equality law to AI deployments.

Why it matters Operationalizes AI governance into nine discrete risk categories with enforceable expectations, shifting compliance from principle-based statements to measurable outcome-based standards.

The UK's AI Risk Management Toolkit creates a structured nine-category framework for assessing AI systems in operation: legal/regulatory compliance, fairness, transparency, accountability, technical robustness, security, and related dimensions. By explicitly integrating existing legal obligations (data protection, equality law) into AI governance, the toolkit signals AI compliance is not a separate domain but an integration point for established regulatory frameworks.

For IT and infrastructure teams, this matters because it operationalizes responsible AI language into a reference model regulators will use to evaluate organizational readiness. Organizations deploying AI in regulated sectors should treat these nine categories as a baseline expectation for governance documentation. The toolkit positions itself as applicable to private sector organizations, indicating future enforcement actions will likely reference these categories as the standard of care for AI deployments.

Read the original at tlt.com ↗
Nothing in today's briefing matches that.

Listening

Podcasts and talks
  • CNCF Blog

    KubeCon + CloudNativeCon North America 2026: Build your SRE journey

    CNCF highlights how SREs should approach KubeCon NA 2026 (Nov 9–12, Salt Lake City) with sessions spanning observability, incident response, autoscaling, and networking. The focus emphasizes learning production patterns and failure stories from peers operating cloud-native systems at scale.

  • CNCF Blog

    KubeCon + CloudNativeCon North America 2026: Build your infrastructure engineer journey

    CNCF outlines infrastructure engineer track for KubeCon NA 2026, emphasizing cluster scaling, multi-cluster networking, GPU/AI workload support, platform engineering, and operational visibility. Sessions address how teams balance reliability, security, and cost while supporting diverse workload types and infrastructure complexity.

Vendor Radar

Last 7 days · arrows compare with the 7 before

Most active

In today's briefing

What changed this week

Last 7 days vs the 7 before

Biggest moves

Trending topics

agent · automation · Agentic AI · OpenTelemetry · observability · MCP · AIOps · SRE · Digital Twin · RAG · LLM · Zero Trust

Get the daily briefing

Today's 3 things that matter and every story with why it matters, in your inbox each morning. Free, and you can unsubscribe at any time. Prefer a reader? Follow the RSS feed.