You're reading the Friday, October 2, 2026 edition. Today's briefing →
Toronto
Friday, October 2, 2026No. 110

Digital Plumber

Plumbing the information age

AI-curated intelligence for people who run networks. Daily coverage of AIOps, network automation, agentic operations, AI infrastructure, security and the vendors shaping them.

Today's 3 things that matter

Picked by the AI editor
  1. Routing·Vendor release

    Azure ExpressRoute and VPN Gateway fail across 19 regions

    Microsoft Azure experienced a major network outage on September 30 affecting 19 regions, impairing ExpressRoute and VPN Gateway connectivity.

    Why it matters Demonstrates cascading failure risk when control-plane changes align with maintenance windows; enterprises must segregate management and data-plane maintenance schedules.

  2. Telco·Industry news

    Verizon moves RAN operations to agentic AI with human guardrails

    Verizon is transitioning from rule-based RAN automation toward agentic systems that reason across network domains (RAN, transport, others) and take autonomous action on unscripted problems.

    Why it matters Moves automation beyond static scripts toward cross-domain AI reasoning (Level 4/5 network ops); defines guardrails for operator control and auditability critical for production telco environments.

  3. Security·Primary source

    Cisco warns of October advisories for NX-OS, APIC and Meraki

    Cisco PSIRT will publish security advisories on October 7, 2026 for multiple products including NX-OS, Application Policy Infrastructure Controller, and Meraki, with security hardening releases.

    Why it matters Advance notice enables network teams to prepare patches for Cisco infrastructure before disclosure; AI-assisted vulnerability discovery signals accelerating attack surface exposure.

Today's briefing

What happened, and why it matters

12 stories · 9 topics · Updated 7:26 PM ET

OpenTelemetry maps its instrumentation ecosystem from SDKs to Collector

OpenTelemetry · Sep 25, 2026 · Primary source

OpenTelemetry blog post documenting the components and breadth of OTel instrumentation: APIs, SDKs, semantic conventions, the OTLP protocol, and tools like the Collector. Explains how instrumentation actually observes system behavior in production.

Why it matters Network ops engineers choosing OTel for network telemetry need to understand which instrumentation layer (automatic vs. manual, agent vs. collector gateway) fits their NOC observability strategy.

The OpenTelemetry project published a foundational guide to its instrumentation ecosystem as OTel approaches de facto standard status post-CNCF graduation. The post breaks down the layers: the OpenTelemetry API and protocol define how telemetry is created and exchanged (vendor-agnostic), while instrumentation is what actually observes system behavior—either automatically (via hooks into common libraries) or manually (via explicit SDK calls in application code). For network observability in a NOC context, automatic instrumentation matters most: OTel can hook into DNS libraries, HTTP clients, and database connectors to generate spans and metrics without code changes. The Collector acts as the central aggregation point, applying sampling, filtering, and routing logic before exporting to multiple backends. Semantic conventions ensure that a 'request.duration' metric means the same thing across all teams and backends. By documenting the full stack, this post helps NOCs understand why they don't have to choose between OTel vendors—the standard itself is the abstraction layer. Network teams can emit flow data, BGP events, and DNS metrics via OTLP and consume them in Datadog, Dynatrace, or open-source backends simultaneously, as long as they follow the semantic conventions.

Read the original at opentelemetry.io ↗

AWS adds AI investigation for CloudTrail and journald log collection

AWS News Blog · Sep 25, 2026 · Vendor release

AWS announced new observability capabilities: AI-powered investigation for CloudTrail, native journald log collection, new pipeline processors for GeoIP and RDS, and OpenTelemetry-based telemetry expansion across services (HealthOmics, WorkSpaces, MSK). CloudWatch Container Insights now supports fractional GPU scheduling.

Why it matters Network ops teams using AWS for hybrid networks and multi-cloud need visibility into how AWS is embedding OTel natively and expanding CloudTrail/CloudWatch signal collection—key for correlated incident investigation across infrastructure layers.

AWS published its monthly observability roundup covering August and September 2026 updates. CloudTrail now supports AI-powered investigation via Amazon Q, allowing plain-language queries ('who accessed this resource and when?') instead of manual log parsing. Log collection expanded with native journald support (critical for systemd-heavy environments) and new CloudWatch Logs pipeline processors for GeoIP enrichment, RDS event parsing, and XML extraction—all natively on ingest, reducing agent-side processing overhead. Multiple AWS services (HealthOmics, WorkSpaces, MSK, others) began publishing richer telemetry to CloudWatch, much of it built on OpenTelemetry—signaling AWS's pivot toward OTel as the native instrumentation standard rather than CloudWatch-proprietary agents. CloudWatch Container Insights extended support for GPU scheduling granularity (down to 1/8 of an L4 GPU), with metrics exposed through Insights and available via OTLP. For network ops in hybrid AWS/on-prem environments, this means: (1) CloudTrail investigation becomes accessible without custom log parsing, (2) OTel metrics from AWS services can flow into third-party observability platforms via native OTLP export, and (3) GPU metrics now available for AI workload monitoring—relevant for network teams supporting inference and training jobs.

Read the original at aws.amazon.com ↗

102.4T switch silicon ships as 1.6T Ethernet nears

Semiconductor Insight · Oct 2, 2026 · Analysis

Market analysis of 2026 Ethernet switch ASIC landscape covering 800G production deployments, 102.4T silicon (Broadcom Tomahawk 6, Cisco Silicon One G300), and 1.6T-ready platforms. Details 224G SerDes as foundational technology for 1.6T and beyond, with Ethernet Alliance roadmap context.

Why it matters Synthesizes ASIC and optics roadmaps for 2026-2027 fabric planning, confirming 800G as deployment standard and providing visibility into 1.6T timeline and vendor silicon options.

Semiconductor Insight analysis maps the 2026 Ethernet fabric landscape across switch silicon generations: 800G remains the active deployment speed across hyperscalers (NVIDIA Spectrum-4, Broadcom Tomahawk 5, Cisco Silicon One G200); 102.4T silicon is entering production (Broadcom Tomahawk 6 shipping, Cisco Silicon One G300 in early deployments); 1.6T-capable platforms are sampling to qualified customers. The analysis emphasizes 224G electrical I/O as the enabling technology for 1.6T, with IEEE 802.3dj defining 8×200G lane configurations. Broadcom Tomahawk 6-Davisson combines 102.4T switching with co-packaged optics (CPO), while Cisco develops 1.6T and 800G optical connectivity via Silicon One G300 with modular optics. Ethernet Alliance September 2026 roadmap identifies 100G–800G as active speeds, 1.6T as emerging with production-readiness expected 2027, and 3.2T on the 2029+ horizon. For network architects: 800G procurement remains the safe immediate choice; 1.6T pilots should target Q2 2027 timeframe when silicon, optics, and test equipment mature.

Read the original at semiconductorinsight.com ↗

Azure ExpressRoute and VPN Gateway fail across 19 regions

Microsoft Azure Status / BigGo Finance · Sep 30, 2026 · Vendor release

Microsoft Azure experienced a major network outage on September 30 affecting 19 regions, impairing ExpressRoute and VPN Gateway connectivity. Root cause was a recent regional gateway management service change coinciding with unrelated infrastructure OS maintenance.

Why it matters Demonstrates cascading failure risk when control-plane changes align with maintenance windows; enterprises must segregate management and data-plane maintenance schedules.

The outage began at 20:30 UTC on September 30, 2026, affecting connectivity services across Japan West, US regions, Europe, and Asia-Pacific. Customers experienced degraded or completely interrupted ExpressRoute connections and VPN Gateway management failures. Microsoft's root cause analysis identified a recent change to the regional gateway management service that coincided with unrelated infrastructure OS maintenance work on the same systems. This is a textbook example of change coordination failure: two independent operational changes created an unexpected interaction. For network operators running hybrid cloud on Azure, this highlights the importance of: (1) staggered maintenance windows across control and data planes, (2) canary deployments of management service changes, and (3) independent rollback paths for each subsystem. Microsoft rolled back the problematic change and confirmed mitigation by 10:15 p.m. EDT (Sept 30). A detailed post-incident review was scheduled within 14 days.

Read the original at finance.biggo.com ↗

Cisco warns of October advisories for NX-OS, APIC and Meraki

Cisco Security · Sep 30, 2026 · Primary source

Cisco PSIRT will publish security advisories on October 7, 2026 for multiple products including NX-OS, Application Policy Infrastructure Controller, and Meraki, with security hardening releases. Vulnerabilities were discovered using frontier AI models during internal testing.

Why it matters Advance notice enables network teams to prepare patches for Cisco infrastructure before disclosure; AI-assisted vulnerability discovery signals accelerating attack surface exposure.

Cisco announced a seven-day advance notice of upcoming October 7, 2026 security advisories covering NX-OS Software for MDS 9000, Nexus 3000, 7000, and 9000 Series Switches, Nexus 9000 in ACI mode, UCS Fabric Interconnects, Application Policy Infrastructure Controller, Finesse License On-Prem, and Meraki platforms. The advisory does not specify CVE numbers or CVSS scores, as these will be disclosed on publication. However, the security advisory explicitly notes that vulnerabilities were identified through frontier AI models used in conjunction with existing internal security testing processes. This represents an operational shift in how Cisco discovers and validates flaws. For network operations teams managing Cisco infrastructure at scale, this signals a need to accelerate patch windows; the advance notice is intended to allow customers to prepare remediation strategies before disclosure details become public. The timing aligns with Cisco's risk-based disclosure process, which publishes hardening releases on the first and third Wednesday of each month, providing predictability for change windows.

Read the original at sec.cloudapps.cisco.com ↗

Itential catalogs 56 production-ready MCP servers for network stack

NetPilot · Oct 2, 2026 · Analysis

Forward Networks moved Forward AI to GA in April 2026, Arista expanded AVA's agentic framework in Q1, Selector AI pushed agentic multi-domain AIOps into NetOps conversations, and Itential catalogs 56 production-ready MCP servers covering nearly every layer of the network stack. The 2026 ecosystem spans vendor-specific tools, open-source options, and MCP-connected agents for multi-vendor environments.

Why it matters Itential FlowAI exemplifies governance-first agentic workflows across network, cloud, ITSM, and security—the emerging standard for operators who need safe, orchestrated agentic operations rather than point tools.

Tools ranked span Cisco AI Assistant for Catalyst Center native management, Juniper Marvis with digital-twin agents, Arista AVA for EOS operations, Forward Networks for deterministic change validation, and Selector AI for multi-domain AIOps. Analysis signals consolidation around three agent patterns: vendor-native (Cisco, Arista, Juniper), multi-domain orchestrators (Selector, Itential), and open-source frameworks (Aurora, K8sGPT). MCP has emerged as the cross-platform glue—56 production servers now connect agents to DNS, firewalls, load balancers, NetBox, BGP routing intelligence, and cloud platforms. Maturation signals: agentic operations is no longer about individual agents but orchestration, governance, and safe delegation across vendor silos. Itential's guide documents official Cisco Catalyst Center MCP server (Aug 2026), making devices, clients, sites queryable through natural language backed by Cisco DevNet.

Read the original at netpilot.io ↗

Verizon moves RAN operations to agentic AI with human guardrails

RCR Wireless · Oct 1, 2026 · Industry news

Verizon is transitioning from rule-based RAN automation toward agentic systems that reason across network domains (RAN, transport, others) and take autonomous action on unscripted problems. The operator retains control: vendors supply domain agents, but only the operator holds network topology and institutional knowledge to arbitrate between them and define intent for consequential changes.

Why it matters Moves automation beyond static scripts toward cross-domain AI reasoning (Level 4/5 network ops); defines guardrails for operator control and auditability critical for production telco environments.

Verizon CTO Anil Guntupalli distinguished between automation ('what we know, scripting around it') and autonomy ('what we don't know, reasoning around it and acting autonomously within domains'). The operator wants AI systems that reason across RAN, transport and other domains and coordinate with one another to solve problems not already scripted by humans. Verizon argues vendors can supply powerful RAN agents, but only the operator has the network topology and institutional knowledge to arbitrate between them. The model maintains human control over intent definition, consequential changes, and major incident oversight. Verizon wants to keep the number of sub-agents limited enough that reasoning chains remain traceable, feeding provenance back into training and improvement. This reflects broader telco shift from siloed automation toward closed-loop, self-optimizing, and increasingly agentic networks including AI-driven RAN and core operations. For ops teams, the key technical challenge is designing agent boundaries and data flows so operator intent can reliably override or constrain autonomous behavior.

Read the original at rcrwireless.com ↗

Bell Canada and Cisco to build sovereign AI infrastructure

Fierce Network · Sep 29, 2026 · Industry news

Bell Canada and Cisco signed a memorandum of understanding to collaborate on sovereign AI infrastructure for Canada, pairing Bell's data center, network and operations assets with Cisco's AI, security, observability and infrastructure management technologies. The partnership targets Canadian organizations needing to deploy, secure and manage AI infrastructure inside the country with greater control over sensitive workload location.

Why it matters Positions Bell as front-runner in Canada's sovereign AI market through its AI Fabric strategy; enables Canadian telcos to capture AI infrastructure workloads without cross-border data governance friction.

Bell and Cisco will focus on three areas: building a Canadian platform for sovereign AI, supporting modular AI infrastructure deployments and developing flexible commercial models. The collaboration combines Bell's Canadian data centre space, power, cooling, physical security, connectivity and operations services with Cisco's AI infrastructure, security, observability and management technologies, including assessment of Cisco AI PODs that provide modular building blocks for training and inference workloads. Bell and Cisco will explore flexible consumption models intended to let customers better match costs with how they use or reserve AI infrastructure. Bell recently launched a cybersecurity AI model in partnership with Cohere, and earlier this month detailed plans to quadruple capacity of its planned AI data centre hub in Saskatchewan. The first portion of Bell's 300-megawatt AI Fabric Data Centre in Sherwood, Saskatchewan, dubbed a 'data hall,' is expected to become operational in first half of 2027, with total estimated capital investment exceeding $50 billion. For Canadian enterprises and regulated sectors (government, finance, healthcare), this removes friction from multi-cloud strategies and simplifies compliance.

Read the original at fierce-network.com ↗

OpenTelemetry details APIs, protocols and semantic conventions

OpenTelemetry Blog · Sep 25, 2026 · Primary source

OpenTelemetry blog post covering the breadth of the instrumentation ecosystem, including APIs, SDKs, protocols, semantic conventions, and tools like the Collector that enable comprehensive observability across applications and infrastructure.

Why it matters Provides practitioners guidance on choosing and integrating OpenTelemetry instrumentation across heterogeneous environments, essential for unified observability in cloud-native operations.

The post outlines OpenTelemetry's modular architecture: APIs and protocols define how telemetry is created and exchanged, while instrumentation libraries are what actually observe runtime behavior. This distinction matters because operators must understand which components to deploy where. The Collector serves as a central processing and export point, enabling flexibility in how signals reach backends. The post emphasizes semantic conventions as a cross-language standard for naming and structuring observability data, reducing friction when migrating between tools or consolidating signals from polyglot systems. For SREs and NetOps teams managing diverse stacks, this ecosystem view clarifies the decision tree: choose instrumentation appropriate to your language/framework, use the Collector for processing/sampling, and export to compatible backends. The post grounds these concepts in practical scenarios, making it accessible to operations teams evaluating observability platforms.

Read the original at opentelemetry.io ↗

NANOG 98 agenda features AI agents and digital twins

NANOG News · Sep 25, 2026 · Primary source

NANOG 98 conference agenda published with sessions covering global connectivity infrastructure, AI agent deployment in network operations, digital twins, and enterprise-to-cloud networking challenges, keynotes from Paul Brodsky (TeleGeography) and John Evans (infrastructure automation).

Why it matters Practitioners can preview peer presentations on production deployment of automation and AI agents in networks, digital twin implementations, and international routing architecture—directly relevant to operational challenges.

The NANOG 98 agenda, released September 25, 2026, reflects shifting priorities in network operations. The spotlight theme of 'Global Connectivity' maps to real operational concerns: routing geography shifts, peering strategies for latency-sensitive traffic, and managing complexity across multi-region deployments. Sessions on agent engineering in network operations are particularly relevant—speakers address guardrails, tool selection, failure handling, and human approval workflows, moving beyond theoretical LLM capability to practical deployment constraints. Digital twin sessions (e.g., 'What Building a Network Emulation Platform Taught Me About Digital Twins' and 'Your Change Deserves a Dress Rehearsal') indicate maturation of testing methodologies for high-risk changes. The keynote from John Evans (founder roles at Cisco, AWS, BT, Cariden) promises insider perspective on why autonomous networks underperform self-driving cars—a reality check on current capabilities. Multi-path traffic engineering workshops address concrete problems in hyperscale networks. For senior network engineers and automation architects, this agenda previews peer solutions to problems they're likely facing: how to safely deploy intent-driven systems, validate changes at scale, and navigate international interconnection agreements. Registration open through October 17.

Read the original at nanog.org ↗

OpenAI releases GPT-6.1 Sol at one-fifth of Astra pricing

Unite.ai · Sep 29, 2026 · Industry news

GPT-6.1 Sol nearly matches GPT-6 Astra's intelligence on agentic coding, computer use, and professional work at one-fifth of Astra's token prices. Released at $2 per million input tokens and $10 per million output tokens.

Why it matters GPT-6.1 Sol delivers Astra-tier capability at 80% lower cost, reshaping inference cost calculations for developers optimizing agentic coding and multi-model deployment strategies.

OpenAI announced GPT-6.1 Sol on September 29, 2026 at DevDay, expanding the GPT-6 family with a cost-optimized variant. The model achieves near-parity with Astra on agentic coding, computer use, and professional tasks while costing one-fifth as much on standard token rates ($2/$10 vs $10/$50). Available to Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex via gpt-6.1-sol. OpenAI introduced Ultrafast, a premium speed tier offering up to 8x faster token generation (300 tokens/sec) in Codex and 6x faster in API, with GPT-6 Astra Ultrafast available now and GPT-6.1 Sol variant coming soon. For infrastructure teams, the pricing structure matters: at $2/$10 per million tokens, GPT-6.1 Sol undercuts Claude Opus 5.5 (released same week at $4/$20) by 50% while matching DeepSWE coding benchmarks, fundamentally altering cost-capability tradeoffs for production deployments.

Read the original at unite.ai ↗

BIS finds 55% of AI funding comes from AI firms

Crypto Briefing · Oct 2, 2026 · Industry news

A Bank for International Settlements study found that 55.2% of incoming investment into AI firms between 2021 and 2025 came from other AI firms. The research, published as BIS Bulletin 137 on October 1, 2026, maps a tightly looped financing web where suppliers often bankroll their own customers.

Why it matters Most sectors do not raise the majority of their outside capital from their own peers, signaling unusual capital concentration and ecosystem dynamics critical for understanding AI market structure and potential systemic risks.

The BIS study reveals a structural anomaly in AI financing: more than half the capital flowing into AI companies originates within the sector itself, creating circular funding patterns where suppliers and customers fund each other. This closed-loop structure stands in stark contrast to traditional industries, where external capital sources typically dominate. The finding carries implications for investors assessing sector health and regulators evaluating competitive dynamics and market concentration. While the research identifies both benefits and risks in these circular relationships among AI firms, the fundamental concern is how this inward-looking capital flow affects pricing dynamics, competitive entry barriers, and overall market stability. Media coverage on October 2 zeroed in on the structural uniqueness of the arrangement. This matters for practitioners evaluating AI vendor ecosystems and supply chain dependencies: understand that your software stack may exist within a self-reinforcing funding loop that could shift rapidly if external capital dries up or consolidates further.

Read the original at cryptobriefing.com ↗
Nothing in today's briefing matches that.

Listening

Podcasts and talks
  • Packet Pushers - Network Automation Nerds

    NAN132: The AI-Augmented Engineer

    Packet Pushers Network Automation Nerds podcast episode exploring how AI augments network engineers in automation workflows. The show examines practical approaches to integrating AI assistance into daily engineering tasks without replacing core technical skills.

Vendor Radar

Last 7 days · arrows compare with the 7 before

What changed this week

Last 7 days vs the 7 before

Biggest moves

Trending topics

agent · automation · observability · Agentic AI · OpenTelemetry · MCP · LLM · SRE · Digital Twin · RAG · AIOps · Zero Trust

Get the daily briefing

Today's 3 things that matter and every story with why it matters, in your inbox each morning. Free, and you can unsubscribe at any time. Prefer a reader? Follow the RSS feed.