You're reading the Thursday, October 8, 2026 edition. Today's briefing →
Toronto
Thursday, October 8, 2026No. 116

Digital Plumber

Plumbing the information age

AI-curated intelligence for people who run networks. Daily coverage of AIOps, network automation, agentic operations, AI infrastructure, security and the vendors shaping them.

Today's 3 things that matter

Picked by the AI editor
  1. DC networking·Primary source

    Arista launches 1.6T rack-scale AI switches with liquid cooling

    Arista unveiled a rack-scale AI networking portfolio increasing capacity from 1.6 Pb/s to 6.5 Pb/s with 50% reduction in data center footprint.

    Why it matters Air-cooled 1.6T systems ship Q4 2026 with liquid-cooled variants in Q1 2027, addressing critical AI fabric density constraints while supporting Linear Pluggable Optics to reduce power consumption by 60%.

  2. DC networking·Primary source

    Broadcom pushes scale-up, scale-out and scale-across AI networking

    Broadcom announced innovations in scale-up, scale-out, and scale-across AI networking at the 2026 OCP Global Summit, including Tomahawk 6, Tomahawk Ultra, and Jericho 4 Ethernet switches, Thor Ultra 800G and 400G NICs, and third-generation TH6-Davisson Co-Packaged Optics designed for frontier AI models requiring exponential bandwidth.

    Why it matters The portfolio addresses bandwidth and power efficiency requirements for AI clusters exceeding 100,000 accelerators through native 1.6T ports, 224G SerDes, and 65% optics power reduction versus pluggable modules.

  3. Research·Primary source

    OpenTelemetry Java agent previews 3.0 with new telemetry defaults

    OpenTelemetry Java agent 2.32.0 RC previews 3.0 release targeted for October 2026, introducing changes to telemetry conventions, semantic attributes, and instrumentation defaults.

    Why it matters Teams using Java agent must test against observability dashboards and pipelines before 3.0 becomes default; semantic convention changes affect metric/span naming across deployments.

Today's briefing

What happened, and why it matters

15 stories · 8 topics · Updated 1:26 PM ET

Nautobot and NetBox diverge as source-of-truth platforms

NetPilot · Oct 2, 2026 · Analysis

Comprehensive comparison of Nautobot and NetBox as network source-of-truth platforms, analyzing architectural differences, extensibility, automation capabilities, and ecosystem maturity. NetBox Labs backing vs Network to Code stewardship shapes divergent evolution paths.

Why it matters Practitioners choosing a network source of truth need concrete guidance on automation-first (Nautobot) vs documentation-first (NetBox) trade-offs and AI/MCP server support.

The article provides a detailed feature matrix comparing Nautobot v3.0.2 (Dec 2025) and NetBox v4.4.8 (Dec 2025) across core dimensions: lineage, licensing, extensibility models, jobs/automation engines, GraphQL support, and ecosystem maturity. Nautobot emphasizes first-class automation with native Jobs engine, scheduling, approvals, and Kubernetes-native execution; NetBox offers background bulk operations and custom scripts. Both support REST and GraphQL APIs with SDK support. Key architectural difference: Nautobot treats Apps as first-class citizens with Git as a data source, while NetBox uses a traditional plugin framework. NetBox Labs (commercial steward since 2023) ships NetBox Cloud and Enterprise; Network to Code maintains Nautobot with managed options. NetBox Copilot reached GA in February 2026; Nautobot has NautobotAI suite. MCP servers available from both: NetBox official (read-only, write via paid platform) and Network to Code official plus community servers. For documentation-first teams wanting industry-default choice, NetBox remains strongest; for automation-first NetDevOps teams, Nautobot's rapid iteration (3.1 April 2026 added commercial apps; 3.2 July 2026 added Reports, Job Chaining, IP Ranges) positions it as the evolution path.

Read the original at netpilot.io ↗

Running AI agents securely at scale needs gateways and sandboxes

Packet Pushers · Oct 7, 2026 · Industry news

Packet Pushers podcast episode exploring operational realities of deploying AI agents at scale, covering specialized AI gateways, semantic routing, sandboxing, and observability practices. Host Ned discusses with Michael Levan (Day Two DevOps episode 315).

Why it matters NetDevOps and SRE teams deploying agentic AI need guidance on secure execution, isolation boundaries, observability instrumentation, and operational guardrails before production rollout.

Episode D2DO315 (Oct 7, 2026, 47 minutes) addresses practitioner concerns as AI agents move from experimental to operational status. The discussion covers specialized AI gateways as control planes for agent routing, semantic routing mechanisms to direct agent reasoning appropriately, sandboxing techniques to prevent lateral movement, and robust observability practices to track agent decisions and actions. Michael Levan from Day Two DevOps brings operational experience; the episode reflects real deployment challenges teams face. Key themes: agents lack operational context without proper infrastructure; executing agents securely requires more than permission systems—it needs execution isolation, semantic understanding of routing policies, and full auditability. This aligns with broader 2026 trend (noted in Itential predictions) where networking shifts from scripts to agent-driven, orchestrated systems requiring MCP, AI agents, and orchestration platforms. Observability practices mirror SRE discipline: structured logging, tracing agent decisions, monitoring action execution outcomes. Episode reflects Packet Pushers' focus on practitioner realities rather than vendor marketing.

Read the original at packetpushers.net ↗

Broadcom pushes scale-up, scale-out and scale-across AI networking

Broadcom press release · Oct 8, 2026 · Primary source

Broadcom announced innovations in scale-up, scale-out, and scale-across AI networking at the 2026 OCP Global Summit, including Tomahawk 6, Tomahawk Ultra, and Jericho 4 Ethernet switches, Thor Ultra 800G and 400G NICs, and third-generation TH6-Davisson Co-Packaged Optics designed for frontier AI models requiring exponential bandwidth.

Why it matters The portfolio addresses bandwidth and power efficiency requirements for AI clusters exceeding 100,000 accelerators through native 1.6T ports, 224G SerDes, and 65% optics power reduction versus pluggable modules.

Broadcom announced its latest innovations for the October 12–15 OCP Global Summit in San Jose, featuring Ethernet switches, NICs, PCIe components and integrated optics to power Open Rack Version 3 (ORV3) solutions. The Tomahawk 6 family delivers 102.4 Tbps switching capacity with native 1.6T per-port throughput using 224G SerDes and PAM4 signaling. The TH6-Davisson co-packaged optics variant integrates eight 6.4Tbps optical engines directly into the switch package, reducing optics power by 65% compared to pluggable modules. The portfolio addresses critical challenges in AI infrastructure scaling by reducing switch tier count, minimizing optical complexity, and lowering total cost of ownership. Tomahawk Ultra provides ultra-low 250ns latency for scale-up fabric connectivity, while Jericho 4 enables secure, lossless fabrics for 1M+ XPU clusters. OCP Foundation CEO highlighted how Broadcom's contributions on ORV3 and wide-rack designs solve bandwidth and efficiency challenges for gigawatt-scale AI data centers.

Read the original at investors.broadcom.com ↗

Arista launches 1.6T rack-scale AI switches with liquid cooling

Street Insider · Oct 7, 2026 · Primary source

Arista unveiled a rack-scale AI networking portfolio increasing capacity from 1.6 Pb/s to 6.5 Pb/s with 50% reduction in data center footprint. The 7060XE7 Series uses Broadcom Tomahawk 6 silicon with 224G SerDes and supports air and liquid cooling for 1.6T deployments.

Why it matters Air-cooled 1.6T systems ship Q4 2026 with liquid-cooled variants in Q1 2027, addressing critical AI fabric density constraints while supporting Linear Pluggable Optics to reduce power consumption by 60%.

Arista emphasized that network infrastructure is now a tightly integrated rack-scale backplane essential to maximizing compute. The 7060XE7 Series delivers multiple configurations: 7060XE7-64PS and 7060XE7-64PRS air-cooled rack switches in 4RU with 64 ports at 1.6T each; 7060XE7-64PRS-RV3-L liquid-cooled 2OU platform; and 7060XE7-128PE providing 128 800G ports. Each system supports Integrated Heat Sink and Riding Heat Sink optics for deployment flexibility. The portfolio integrates 224G SerDes technology with LPO support to reduce interconnect power consumption by approximately 60%. EOS software features include AI-focused congestion management, load balancing, resiliency, and deep diagnostics optimized for collective communication patterns. Arista's AI networking revenue doubled to $1.5B in FY2025 with $3.25B guided for 2026.

Read the original at streetinsider.com ↗

Broadcom starts volume shipments of Davisson co-packaged optics

Data Centre & Network News · Oct 8, 2026 · Industry news

Recent developments including AMD's Enosemi acquisition, Lightmatter's Passage L200 (world's first 3D co-packaged optics product), Broadcom's Tomahawk 6 CPO variant 'Davisson', and Ayar Labs/Wiwynn joint reference design for optically connected rack-scale AI systems demonstrate CPO market acceleration.

Why it matters Volume shipments from Broadcom's Tomahawk 6 Davisson CPO variant represent the industry's first production-scale CPO deployment, unblocking widespread adoption of co-packaged optics in hyperscale AI infrastructure.

Co-packaged optics offers clear operating cost advantages primarily through reduced power consumption and improved conversion efficiency. According to Broadcom data, pluggable optical modules consume 15–20 pJ/bit, while CPO systems reduce this by over 50% to 5–10 pJ/bit. The Broadcom Bailly 51.2Tbps CPO switch integrates eight 6.4Tbps Silicon Photonics Optical Engines, with each supporting 64 Tx/Rx channels, totaling 128 400Gbps FR4 optical interfaces connected via standard single-mode fiber. The CPO optoelectronic package mounts directly on the Switch Main Board, providing power and control functions. This tight integration addresses a fundamental scaling challenge where, between 2010 and 2022, global data center network switching bandwidth surged 80 times while optical module power consumption rose 26 times—a mismatch CPO is engineered to resolve. Cisco has demoed CPO prototypes but has not announced commercial offerings, likely due to reliability concerns, whereas Broadcom's field-proven Davisson system is now shipping to early access customers.

Read the original at datacenterdynamics.com ↗

Ethernet switching hits $18.9 billion a quarter on AI demand

The CODEW · Oct 1, 2026 · Analysis

IDC data shows quarterly Ethernet switching at $18.9 billion growing above 40% YoY, with 800G ports comprising two-fifths of data-center switch revenue. Arista, NVIDIA, Cisco, and HPE Juniper are driving AI networking adoption across scale-up and scale-out architectures.

Why it matters Arista and HPE are pushing Ethernet into the rack—a domain NVIDIA's NVLink has held—while optics suppliers race toward 1.6T and co-packaged designs, reshaping network architecture choices for AI operators.

IEEE P802.3dj, which includes 1.6Tb/s Ethernet specifications, is advancing with the Ethernet Alliance's Technology Exploration Forum running October 7–8 in Mountain View. Market adoption shows 800G as the 2026 workhorse with 41.2% share of data-center switch revenue. 1.6T is entering early commercial rollout at hyperscalers mid-2026, with the standard pattern being a 1.6T switch tier feeding 800G NICs through 2×800G breakout. Broadcom Tomahawk 6 systems and Arista 7060XE7 present native 1.6T OSFP ports—DR8 and 2×FR4 running as single 1.6T links, with LPO as a host-supported option. The economics shift reflects 800G reaching volume at $40–50K per switch versus 1.6T spines offering flatter topologies with fewer optical links but higher per-port cost. The quarterly data-center Ethernet switching market of $18.9 billion represents continued displacement of InfiniBand at hyperscalers, with Arista alone guiding $3.5 billion in AI fabric revenue for 2026.

Read the original at thecodew.com ↗

Cloudflare redesigns Radar and widens its measurement base

Cloudflare Blog · Oct 8, 2026 · Primary source

Cloudflare redesigned Radar to make real-time global traffic and outage data accessible to broader audiences while preserving technical depth through interactive maps and standardized components. Expanded measurement scale by integrating background telemetry from Cloudflare Challenge Pages.

Why it matters Improves visibility into internet routing incidents and outages for practitioners who rely on real-time operational data; expanded measurement methodology enhances detection accuracy.

Cloudflare's redesigned Radar balances accessibility with technical rigor for network operations practitioners tracking global routing events and outages. The interface introduces interactive mapping and standardized UI components to simplify navigation while retaining detailed BGP, DNS, and traffic metrics. A key technical enhancement expands measurement scale by leveraging background telemetry collected from Cloudflare Challenge Pages—when users encounter Cloudflare error pages, their browsers run lightweight measurement fetches to multiple CDN providers including Cloudflare, Amazon CloudFront, Google, Fastly, and Akamai, generating real-user measurement data at scale while maintaining privacy. This methodology now ranks Cloudflare as fastest across 74% of top 1,000 global networks by connection time (TCP handshake completion). The connection-time metric directly correlates to perceived user experience and serves as a proxy for network performance quality, enabling operators to identify infrastructure or routing issues affecting specific AS paths or geographies.

Read the original at blog.cloudflare.com ↗

SRE agents move from experiments to standard operations

DEV Community · Oct 7, 2026 · Industry news

Comprehensive overview of ReAct pattern, Azure SRE Agent, PagerDuty's multi-agent approach, and Datadog Bits AI SRE, with concrete examples of gradual autonomy stages and operational metrics. Shows how SRE agents have moved from experimental to operational standard.

Why it matters Practitioners need to understand ReAct pattern dominance and autonomy staging models to architect agent systems that balance safety with efficiency in production operations.

The article details how AI agents in SRE and DevOps have established the ReAct (Reason and Act) pattern as the dominant architectural approach for operations agents in 2026. Azure SRE Agent's Deep Context function connects code repositories, logs, incidents, and Azure resources into a single context graph with persistent memory. PagerDuty's multi-agent approach splits responsibilities: SRE Agent for autonomous incident analysis and fixes, Scribe Agent for transcribing Slack/Zoom conversations, and Shift Agent for schedule conflict resolution, delivering 50% faster incident resolution. Datadog Bits AI SRE operates as a 24/7 teammate that continuously maps service dependencies and deployment histories. The piece emphasizes gradual autonomy as enterprise standard: Stage 1 read-only, Stage 2 low-risk operations (cache clearing, read-replica scaling), Stage 3 standardized fixes with time-window approval, Stage 4 high-risk operations disabled by default. Microsoft saves 20,000 engineering hours per month with 1,300 Azure SRE agents, demonstrating concrete ROI for operations teams that stage autonomy carefully.

Read the original at dev.to ↗

Netskope opens agent marketplace for SOC and NOC teams

Globe Newswire · Oct 6, 2026 · Vendor release

Netskope announced AgentSkope Marketplace providing security and network teams a single console to browse, deploy, and manage agents with unified capacity-credit pooling instead of per-agent procurement.

Why it matters NetOps and SecOps teams can now treat agents as shared operational resources with unified capacity management, enabling faster scaling and reduced administrative overhead.

Netskope's AgentSkope Marketplace consolidates agent management through a single catalog and capacity-credit pool, eliminating per-agent procurement friction. The platform was designed specifically for agentic SOC and agentic NOC use cases—teams deploy managed agents within Netskope One that execute end-to-end security and networking workflows, all constrained by existing policies, controls, guardrails, and compliance frameworks. This matters operationally because teams previously had to evaluate each agent independently, provision separate resources, and manage different control planes. Now, teams allocate shared capacity credits and scale agents based on actual usage insights rather than peak-load estimates, reducing administrative overhead and enabling rapid experimentation with new agents.

Read the original at globenewswire.com ↗

Trustgrid MCP server lets agents run live node diagnostics

Trustgrid · Oct 6, 2026 · Primary source

Trustgrid released an MCP Server minor release enabling agents to run interactive tools requiring WebSocket connections, such as ping and arping, directly on network nodes for real-time diagnostics.

Why it matters NetOps agents can now execute interactive network diagnostics in real time through MCP, closing the gap where agents were limited to read-only or async operations.

Trustgrid's October 6 release extends its MCP Server to support interactive tools requiring WebSocket connections—specifically ping and arping operations on network nodes. This addresses a critical operational limitation: previous agent implementations were constrained to synchronous API calls or asynchronous operations, making real-time interactive diagnostics difficult. With WebSocket support, agents can now execute live network diagnostics with full bidirectional communication, enabling faster root-cause analysis when latency or reachability issues occur. This fits naturally into the MCP architecture, where agents interact with network tools through standardized protocol interfaces, reducing the need for custom integrations.

Read the original at docs.trustgrid.io ↗

Agent-to-agent MCP links create new trust boundary risks

TechnoSports · Oct 6, 2026 · Analysis

Security researchers warn that agent-to-agent MCP communications introduce trust boundary risks; unauthorized data access can occur when agents exchange instructions across local system boundaries if input validation is insufficient.

Why it matters Operations teams deploying multi-agent systems via MCP must implement strict access controls and input validation at every agent hop, treating agent-to-agent traffic with the same security rigor as external APIs.

Security research published October 6, 2026, warns that agent-to-agent MCP communications introduce trust boundaries that organizations often overlook. Unauthorized data access risks emerge in agent-to-agent MCP implementations, particularly where agents exchange forwarded instructions across local system boundaries. The research emphasizes that developers can no longer assume internal network traffic is inherently benign simply because requests originate from approved software nodes. The security posture must shift to rigorous input validation at every hop, ensuring forwarded instructions undergo the same scrutiny as initial user prompts. This demands strict access control lists and comprehensive monitoring telemetry—no longer optional when automating complex multi-step workflows where one compromised agent could laterally access resources intended for other agents.

Read the original at technosports.co.in ↗

Bell Saskatchewan AI data centre targets first hall in 2027

iPhone in Canada · Oct 1, 2026 · Industry news

Bell reported that more than 500 people are working at the Rural Municipality of Sherwood site for its Saskatchewan AI data centre, with the first data hall expected to come online in H1 2027. The proposed expansion would require an additional 900 MW of power using dedicated natural gas generation funded by Bell.

Why it matters Bell and Cisco have signed an MOU to collaborate on sovereign AI infrastructure in Canada, keeping AI systems and data under Canadian control as telcos pivot from connectivity to AI infrastructure services.

Bell announced progress on its Saskatchewan AI data centre with 500+ workers now on site at the Rural Municipality of Sherwood location, involving over 30 local contractors and suppliers with most workforce from Saskatchewan. The first data hall is still expected to come online in the first half of 2027. This build-out is part of Bell's broader AI Fabric strategy launched earlier in 2025 with initial hydro-powered facilities in British Columbia.

The Saskatchewan phase represents Bell's expansion of AI infrastructure capacity eastward. The expansion would be four times the size of current construction and rely on dedicated natural gas generation funded by Bell, though actual build depends on customer commitments, commercial agreements, permits, and approvals. Bell positions itself as a full-service AI provider offering hardware infrastructure, AI strategy, and application development for Canadian enterprises and governments—a strategic pivot from pure connectivity toward infrastructure-as-a-service in the AI economy.

Read the original at iphoneincanada.ca ↗

OpenTelemetry Java agent previews 3.0 with new telemetry defaults

Grafana Labs / OpenTelemetry Blog · Oct 6, 2026 · Primary source

OpenTelemetry Java agent 2.32.0 RC previews 3.0 release targeted for October 2026, introducing changes to telemetry conventions, semantic attributes, and instrumentation defaults. Database and code conventions stabilize; messaging adopts experimental conventions; Hibernate, Hystrix, Twilio instrumentation default to off.

Why it matters Teams using Java agent must test against observability dashboards and pipelines before 3.0 becomes default; semantic convention changes affect metric/span naming across deployments.

The 2.32.0 release serves as a preview and release candidate for OpenTelemetry Java agent 3.0, enabling practitioners to test breaking changes before the final October 2026 release. Key changes include stabilization of database and code semantic conventions, adoption of newer (still experimental) messaging conventions that alter span parent-child relationships, and several instrumentation defaults changing—Hibernate, Hystrix, and Twilio instrumentation now require explicit re-enablement. The release also removes the Zipkin exporter and enables invokedynamic instrumentation by default. This is critical for SRE and platform teams managing distributed tracing pipelines, as metric names, span attributes, and cardinality may shift. The OpenTelemetry specification v1.61.0 was released September 14, 2026, and the JavaScript SDK 3.0 is scheduled for October 15, 2026, creating a coordinated multi-language update window. Teams should test this RC against production-like dashboards and alert rules to validate compatibility before 3.0 becomes mandatory.

Read the original at opentelemetry.io ↗

Mistral Large 4 tops cybersecurity benchmark at 82 percent

Mistral AI · Oct 6, 2026 · Primary source

Mistral AI released Mistral Large 4 on October 6, 2026 as a public preview, a 1.05-trillion-parameter multimodal model with 52 billion active per token and a 1.6-billion-parameter vision encoder, with open weights promised by end of month. On cybersecurity benchmarks it scores 82% on Artificial Analysis Cyber Index (reproduce-and-patch real vulnerabilities), which Mistral claims is the highest of any model.

Why it matters The 82% reproduce-and-patch score and 93% on Cybench make Large 4 a critical option for security teams needing a model willing to engage adversarial material under controlled conditions, especially with open weights and self-hosting capability.

Mistral Large 4 (nicknamed "Le Chonk") is Mistral's flagship offering and represents European AI capability at trillion-scale. Architecture: 1.05T total parameters, mixture-of-experts with 52B active per token, 1.6B vision encoder, native multimodal (text+images in, text out). Trained on 3,800 Nvidia Grace Blackwell GPUs in Mistral's European datacenters (~2 months, ~10MW). Context: 1M tokens (524K listed on some preview APIs). Knowledge cutoff: training data unspecified. Pricing: $1.36/M input, $4.18/M output. Benchmarks: DeepSWE 61.7% (agentic repo engineering, tied with GLM-5.3), Terminal-Bench 28.3% (long terminal sessions—a weak point), SWE-Atlas-QnA 59.4%, Legal (Harvey's benchmark) 15.83% (beating GLM-5.3 8.33%, GPT-6 Astra 5.42%), Finance Agent v2 54.7% (behind GLM-5.3 55.8% but ahead of Astra 53.5%). Cybersecurity: Artificial Analysis CyberGym-E2E 81.7% (first place), Cybench 93%, B3 attack resistance 93.3%. Artificial Analysis Intelligence Index: 38 (tied with GPT-6 Luna, highest Western open-weights model on that chart). Weights release: promised by October 27. License terms not yet public. Key caveat: expanded cyber capabilities in red-team testing may not be available to standard API users.

Read the original at mistral.ai ↗

O'Reilly grounds enterprise AI agents in cited technical sources

AIwire (HPC Wire) · Oct 7, 2026 · Industry news

O'Reilly announced Expert Intelligence, a platform that grounds enterprise AI agents in its repository of practitioner knowledge and technical documentation, with all outputs verified and cited to named authors. The suite provides job-specific agent skills for actionable workplace guidance.

Why it matters This product addresses enterprise AI governance gaps by ensuring AI outputs are auditable and traced to verified sources, enabling compliance with EU AI Act requirements and NIST frameworks where black-box AI carries unacceptable regulatory risk.

O'Reilly launched Expert Intelligence, grounding enterprise AI agents in its corpus of technical books and practitioner documentation with mandatory source citations. This directly addresses a critical governance challenge: traditional LLMs generate hallucinated information; Expert Intelligence constrains agent outputs to verified content and requires citations to named authors and publications. This architectural approach enables IT and compliance teams to audit agent decisions, validate accuracy before deployment, and demonstrate due diligence to regulators under EU AI Act high-risk system requirements and NIST AI RMF frameworks. The shift toward citation-verified AI reflects an emerging market segment of governance-first enterprise AI platforms that treat verifiability as a core requirement rather than an afterthought, particularly important for enterprises managing regulatory liability and audit compliance.

Read the original at hpcwire.com ↗
Nothing in today's briefing matches that.

Listening

Podcasts and talks
  • ipSpace.net blog

    Please Respond: State of Network Automation Survey

    The Network Automation Forum is running the annual State of Network Automation Survey, building on 2025 results to benchmark adoption, tooling gaps, and adoption barriers across the industry.

  • Packet Pushers (IPv6 Buzz)

    IPB209: SREs Are Breaking IPv6 — And They Don't Know It Yet

    Ed and Nick explore why developers and SREs become bottlenecks during IPv6-only transitions, with practical strategies for bridging communication gaps between network and application teams during architectural design.

Vendor Radar

Last 7 days · arrows compare with the 7 before

What changed this week

Last 7 days vs the 7 before

Biggest moves

Trending topics

agent · automation · MCP · Agentic AI · OpenTelemetry · SRE · observability · AIOps · LLM · SASE · inference · Digital Twin

Get the daily briefing

Today's 3 things that matter and every story with why it matters, in your inbox each morning. Free, and you can unsubscribe at any time. Prefer a reader? Follow the RSS feed.