You're reading the Saturday, October 10, 2026 edition. Today's briefing →
Toronto
Saturday, October 10, 2026No. 118

Digital Plumber

Plumbing the information age

AI-curated intelligence for people who run networks. Daily coverage of AIOps, network automation, agentic operations, AI infrastructure, security and the vendors shaping them.

Today's 3 things that matter

Picked by the AI editor
  1. Security·Primary source

    FBI says 86,000 Fortinet firewalls compromised by credential harvesting

    The FBI and Secret Service warned that over 86,000 Fortinet FortiGate firewalls across 194 countries have been compromised through a credential-harvesting campaign.

    Why it matters No CVE exists—attackers use reused and leaked credentials, credential stuffing and password spraying, then crack offline password hashes on GPU clusters, making traditional patch-driven vulnerability management insufficient.

  2. DC networking·Primary source

    Arista scales Ethernet fabric to 144 accelerators per domain

    Arista announced Etherlink SU-144, an Ethernet-based scale-up architecture expanding unified computing domains to 144 accelerators in single-hop and 1,024 XPUs across cross-rack topologies, engineered with AMD, Broadcom, Meta, Microsoft, and Qualcomm via the Ethernet for Scale-up Networks (ESUN) initiative.

    Why it matters Scale-up fabric standardization through ESUN shows industry consensus on Ethernet for tightly coupled GPU/XPU pools, reducing custom engineering overhead and enabling faster cluster deployment for large LLM training and inference workloads.

  3. Security·Primary source

    Cisco and SonicWall patch CVSS 10 authentication bypasses

    Two of the most widely deployed names in enterprise networking issued emergency patches for the same class of bug—Cisco pushed a fix for a critical authentication-bypass flaw in Catalyst SD-WAN Manager and SonicWall for CVE-2026-102255 in SMA 1000.

    Why it matters VPN gateways and SD-WAN controllers are built to be reachable from the public internet, sit in front of everything else on the network, and patching them requires a maintenance window that many IT teams schedule weeks or months out rather than applying same-day, creating a critical priority-patching gap.

Today's briefing

What happened, and why it matters

14 stories · 7 topics · Updated 11:55 AM ET

Operators tackle observability across GPU fabrics of ten thousand

Packet Pushers · Oct 9, 2026 · Primary source

Scott Robohn hosts Christian Adell and David Flores to discuss observability and automation challenges in AI data centers where thousands or tens of thousands of GPUs demand unified visibility. The episode addresses how network operations practitioners handle observability at scale for AI infrastructure.

Why it matters Critical for NetOps/SRE teams managing explosive AI data center growth; reveals operational gaps between observability and automation tooling.

This episode directly addresses the intersection of observability, automation, and networking in AI data centers—a growing operational challenge as organizations scale GPU clusters from thousands to tens of thousands of devices. With AI data centers linking massive GPU pools, observability becomes a fundamental blocker to reliable automation. The discussion explores how practitioners are solving this problem in production, including tooling choices, architectural patterns, and the relationship between telemetry collection and automation reliability. For NetDevOps and SRE practitioners, this captures real production insights into why traditional network automation tools struggle at AI infrastructure scale and what observability-first automation looks like in practice.

Read the original at packetpushers.net ↗

Broadcom guide urges consolidating network observability tools for AI

Broadcom · Oct 5, 2026 · Primary source

Modern hybrid networks spanning multi-cloud, SD-WAN, and unowned internet paths require unified observability that consolidates tool sprawl and bridges visibility gaps. The guide emphasizes unified continuous active path testing, device telemetry, and self-adjusting baselines into a single operational view for accelerated root-cause isolation and AI-ready telemetry.

Why it matters Organizations must consolidate tool sprawl, bridge visibility across unowned cloud and internet paths with active synthetic probes, and build AI-ready telemetry with correlated hop-by-hop evidence to validate automated recommendations before execution.

Modern hybrid networks require unified observability addressing visibility gaps across multi-cloud, SD-WAN and internet paths through continuous active testing and device telemetry in a single operational view. The 2026 industry trend is unified observability with user experience as the primary metric and infrastructure providing diagnostic context. Modern platforms must correlate network and application worlds instantly, providing insights like 'The Zoom call quality dropped because the QoS tag was stripped at the edge router.' For network operations practitioners, outage detection now requires understanding dependencies between user-facing services and underlying network path behavior—a capability demanding integrated synthetic testing, flow analysis, and service telemetry in a single correlated data stream rather than separate tools. This shift eliminates the traditional split between performance monitoring and digital experience monitoring, unifying them into a single correlation layer where every alert carries full network topology and service dependency context.

Read the original at networkobservability.broadcom.com ↗

Broadcom to show Tomahawk 6 and co-packaged optics at OCP

AIwire (HPC Wire) · Oct 8, 2026 · Vendor release

Broadcom announced it will present its latest innovations in scale-up, scale-out, and scale-across AI networking at the OCP Global Summit (Oct 12-15 San Jose), featuring Tomahawk 6, Thor Ultra 800G and Thor 2 400G NICs, and third-generation TH6-Davisson co-packaged optics delivering 102.4 Tb/s with integrated optics for Open Rack v3.

Why it matters CPO maturation and 102.4T switches signal the shift from pluggable optics to integrated architectures, reducing power and improving reliability for multi-million GPU deployments—critical for practitioners sizing AI cluster interconnects.

Broadcom will showcase Ethernet switches, NICs, PCIe components and integrated optics at the October 12-15 OCP Global Summit in San Jose, featuring Tomahawk 6, Tomahawk Ultra, and Jericho 4 Ethernet switches, Thor Ultra 800G and Thor 2 400G AI Ethernet NICs, and the third-generation TH6-Davisson Co-Packaged Optics (CPO) portfolio. The TH6-Davisson delivers approximately 70% reduction in optical interconnect power consumption relative to traditional pluggables—more than 3.5 times lower, resulting in significant TCO and sustainability wins for hyperscale and AI data centers. Improved link stability eliminates manufacturing and test variability typical of pluggable transceivers, reducing flap events which is pivotal for long-duration AI training runs. CPO adoption remains supply-constrained but is no longer research-phase: total shipment volume is expected to range from 10-15k units in 2026, with CPO switches now in production from NVIDIA, Broadcom, and others. For network architects, this marks the transition point where CPO payoff calculations move from theoretical to deployment-ready, especially for high-density scale-out clusters above 10k GPUs where power and link stability directly impact training economics.

Read the original at hpcwire.com ↗

Arista scales Ethernet fabric to 144 accelerators per domain

Arista Networks · Oct 7, 2026 · Primary source

Arista announced Etherlink SU-144, an Ethernet-based scale-up architecture expanding unified computing domains to 144 accelerators in single-hop and 1,024 XPUs across cross-rack topologies, engineered with AMD, Broadcom, Meta, Microsoft, and Qualcomm via the Ethernet for Scale-up Networks (ESUN) initiative.

Why it matters Scale-up fabric standardization through ESUN shows industry consensus on Ethernet for tightly coupled GPU/XPU pools, reducing custom engineering overhead and enabling faster cluster deployment for large LLM training and inference workloads.

Arista is introducing Etherlink SU-144, an Ethernet-based scale-up architecture that expands unified computing domains up to 144 accelerators (XPUs) in a single-hop network and scales to 1,024 XPUs through cross-rack topologies. This massive densification translates to nearly a 50% reduction in the overall data center footprint, resulting in significant time and financial savings. This effort is engineered in partnership with AMD, Arm, Broadcom, d-Matrix, Meta, Microsoft, and Qualcomm through the Ethernet for Scale-up Networks (ESUN) initiative. Scale-up networking—where GPUs share memory and ultra-low latency is required—was historically InfiniBand territory. Arista's multi-vendor ESUN backing signals that Ethernet with RoCEv2 has crossed the threshold from scale-out novelty to scale-up standard, which materially impacts both fabric vendor selection and network engineer skill requirements going forward.

Read the original at arista.com ↗

Broadcom pitches scale-up and scale-across AI networking at OCP

Globe Newswire · Oct 8, 2026 · Vendor release

Broadcom announced it will present innovations spanning scale-up, scale-out, and scale-across AI networking at OCP Summit, featuring Ethernet switches, NICs, PCIe components and integrated optics for Open Rack Version 3 solutions, including Tomahawk 6, Thor NICs, and TH6-Davisson CPO platforms.

Why it matters OCP Summit timing and rack-integrated optics announcements indicate CPO and high-radix switching are transitioning from vendor-specific to open-ecosystem standards, reducing hyperscaler dependency and enabling broader fabric deployment patterns.

Broadcom announced it will present innovations in scale-up, scale-out, and scale-across AI networking at the 2026 OCP Global Summit running October 12-15 in San Jose, featuring Ethernet switches, NICs, PCIe components and integrated optics to power Open Rack Version 3 (ORV3) solutions for scaling AI infrastructure. Broadcom's contributions to OCP include ORV3 rack-based designs and advances in ORW (wide) racks and new wide co-packaged optics that reflect collaboration solving the bandwidth and efficiency challenges facing AI data centers at scale. The emphasis on Open Rack integration signals that CPO is no longer a point product but a foundational building block for standardized AI cluster design. This is particularly relevant for practitioners building or upgrading multi-vendor fabrics, as ORV3 CPO specs establish interoperability boundaries and eliminate custom integration work that previously locked customers into single vendors.

Read the original at globenewswire.com ↗

FBI says 86,000 Fortinet firewalls compromised by credential harvesting

DIESEC / FBI & Secret Service Advisory · Oct 6, 2026 · Primary source

The FBI and Secret Service warned that over 86,000 Fortinet FortiGate firewalls across 194 countries have been compromised through a credential-harvesting campaign. The operation, known as FortiBleed, exploits reused passwords, leaked logins, and outdated password storage to break in, then deletes or changes administrator accounts to seize control.

Why it matters No CVE exists—attackers use reused and leaked credentials, credential stuffing and password spraying, then crack offline password hashes on GPU clusters, making traditional patch-driven vulnerability management insufficient.

More than 86,644 Fortinet FortiGate firewalls across 194 countries have been compromised in a credential-harvesting campaign so aggressive that some victim organizations found themselves locked entirely out of their own systems. Hackers gain access to exposed endpoints by using previously leaked credentials or logins obtained from infostealer logs, credential stuffing, and password spraying attacks, then extract additional authentication data from compromised devices and use a distributed GPU cluster running Hashcat and Hashtopolis to crack offline the stolen password hashes. Attackers create admin accounts that were never on the device and sometimes delete or change the original ones, so legitimate owners lose access, and the intrusion chain has been seen as an entry point for ransomware affiliates linked to INC/Lynx and Payload. Federal authorities are urging organizations to terminate active administrative and VPN sessions, reset credentials, enforce phishing-resistant multi-factor authentication, and restrict management access using trusted host lists. This represents a credential-first attack strategy that bypasses patch management entirely.

Read the original at diesec.com ↗

Cisco and SonicWall patch CVSS 10 authentication bypasses

Tech Insider / CISA KEV · Oct 10, 2026 · Primary source

Two of the most widely deployed names in enterprise networking issued emergency patches for the same class of bug—Cisco pushed a fix for a critical authentication-bypass flaw in Catalyst SD-WAN Manager and SonicWall for CVE-2026-102255 in SMA 1000. SonicWall's advisory states the flaws affect only the SMA 1000 WorkPlace portal and do not touch the separate SMA 100 series or the SSL-VPN functionality built into SonicWall's firewall line.

Why it matters VPN gateways and SD-WAN controllers are built to be reachable from the public internet, sit in front of everything else on the network, and patching them requires a maintenance window that many IT teams schedule weeks or months out rather than applying same-day, creating a critical priority-patching gap.

Cisco and SonicWall spent the first two weeks of October 2026 issuing emergency patches for authentication-bypass flaws scoring near-max CVSS—Cisco's Catalyst SD-WAN Manager and SonicWall's CVE-2026-102255 in SMA 1000. Both enable unauthenticated remote code execution or full system control from the internet perimeter. Network edge appliances show up often in exploited-vulnerability advisories because they are built to be reachable from the public internet, sit in front of everything else on the network, and patching them requires maintenance windows IT teams schedule weeks or months out; detecting post-exploitation activity—reverse shells, unexpected outbound connections from a VPN appliance—often matters as much as preventing the initial breach. SonicWall's advisory carefully scopes damage to SMA 1000 Workplace portal only, not SMA 100 series or SSL-VPN on FireWalls—a distinction that matters for IT teams trying to triage exposure quickly across a mixed fleet.

Read the original at tech-insider.org ↗

MSSP platform automates resolution of 70 percent of incidents

Security Boulevard · Oct 7, 2026 · Analysis

Detection platforms (aiSIEM, aiXDR, UEBA, NDR) and response (aiSOAR) now share one data lake so playbooks act on correlated, enriched incidents with full context, featuring multi-tenant operations with true multi-tier multi-tenancy running 50+ client tenants from a single console, noise reduction via 4,000+ ML models, and response automation resolving about 70% of incidents without analyst intervention.

Why it matters 50.4% of security leaders believe threat detection and risk triage will benefit most from AI automation, and 53% of organizations plan to adopt AI agents for threat detection, signaling that network threat response is shifting from reactive triage to autonomous detection and containment.

Modern MSSP operations require multi-tenant isolation with per-tenant reporting and white-label options, noise reduction through behavioral baselines and telemetry correlation into incidents rather than events, response automation with 100+ out-of-the-box playbooks and automated containment in under 90 seconds, and GenAI-assisted playbook generation from plain language in about 30 seconds. Cisco's September 2026 Splunk announcements illustrate this direction, expanding AI-powered security capabilities across network, cloud, application, and identity data with specialized agents supporting threat hunting, investigation, detection engineering, response, and governance. A Palo Alto Networks case study showed analysts now spend 70% of their time on proactive threat hunting rather than reactive triage after implementing AI automation, reshaping SOC staffing models and shifting network threat operations toward higher-value investigation work rather than alert fatigue management.

Read the original at securityboulevard.com ↗

Cisco puts Anthropic managed agents inside Webex spaces

Startup Fortune / SiliconANGLE / Computerworld · Oct 7, 2026 · Industry news

Cisco announced Claude Managed Agents (cloud-hosted agents from Anthropic) joining Webex spaces and calls for multi-step background work; agents can analyze data, generate content, and coordinate across approved agents within shared team context using MCP integration.

Why it matters MCP integration into Webex and governance patterns Cisco built show how infrastructure ops teams can embed AI agents into collaboration workflows with model choice flexibility and IT-controlled scoping of agent actions.

At WebexOne 2026 (Oct 7), Cisco unveiled Claude Managed Agents as persistent participants in Webex spaces, not just chatbots answering questions. The system allows multi-step workflows: an analytics agent pulls usage numbers when @mentioned, then a presentation agent drops those figures into PowerPoint live, all within the Webex conversation context. The key operational insight is governance: Cisco implemented layered governance so IT can control which agents join which spaces, what tools they can invoke, and how their decisions are audited. The architecture runs on Anthropic's Claude Managed Agents API, with agents operating on separate servers from sandboxes and credentials held in vaults. Cisco also integrated MCP (Model Context Protocol) with OpenAI's dot agents, signaling vendor-neutral tooling. For infrastructure teams, this demonstrates how MCP becomes the connective tissue between agents and operational systems—whether collaboration, observability, or network management. The announcement showed that long-running agents can work overnight and coordinate handoffs between specialized agents without human involvement. Most detailed features ship early 2027, but the architecture is production-ready now.

Read the original at startupfortune.com ↗

Anthropic adds dynamic workflows and vaulted credentials to agents

Anthropic Claude Platform Docs / Releasebot · Oct 9, 2026 · Primary source

Anthropic released dynamic workflows in beta for Claude Managed Agents (Oct 9); workflows enable agents to write and execute multi-agent programs running in background, with phases combining results; new governance: credentials held in vaults, audit trails on all agent actions, two-person approval for sensitive operations.

Why it matters Dynamic workflows and credential management patterns directly translate to infrastructure automation; shows how agents can orchestrate multi-phase tasks (e.g., diagnosis → remediation → validation) without human intervention between steps.

The Oct 9 release adds multiagent_20261001 workflows allowing agents to plan and execute many sub-agents in phases, combining results at end—useful for tasks like document review or incident orchestration. Workflows run in background and tracked via workflow_run events. Operationally, the governance model is designed for infrastructure: credentials never exposed to agent sandbox, stored in separate vault with scoped access; administrative actions logged for audit; two-person approval on Anthropic's side for sensitive operations. Cloud environments with limited networking now scope allowed_hosts to web_search and web_fetch tools, preventing exfiltration. For ops teams building agents on Claude, this release signal is that stateful multi-step workflows with proper isolation and auditability are now table stakes. The Oct 7 release also restricted networking in managed agents by default (allowed_hosts allowlisting), a security stance important for operations environments. Together, these changes position Claude Managed Agents as production-grade for infrastructure use cases where audit and compliance matter.

Read the original at platform.claude.com ↗

Anthropic details scoped credential patterns for agent automations

Claude Blog (claude.dev) · Oct 8, 2026 · Primary source

Anthropic published operational guide for building scheduled Claude Managed Agents automations; core pattern: agents get scoped credentials via MCP servers or shell (not raw API keys), credentials held in vault and swapped at runtime to prevent exposure in sandbox.

Why it matters Practical reference for ops teams designing agent automations; shows how to safely wire agents to infrastructure APIs without embedding credentials, critical for production deployments.

The post walks through six rules for avoiding common agent automation failure modes. Central pattern: agents don't see raw credentials. Instead, MCP servers act as proxies—agents call MCP tools, proxy finds vault credential matching the server's URL and swaps it in as the request leaves sandbox. For shell-based access (e.g., Slack API via curl), agents only see opaque placeholders like $SLACK_BOT_TOKEN; platform swaps real token for allowed hosts. This isolation pattern solves the operational dilemma: agents need real access to infrastructure, but exposing credentials in context is a breach risk. The guide emphasizes credential scoping (one per source), monitoring agent access, and audit trails. For infrastructure ops, this translates directly: if you're building agents to interact with network APIs, IPAM systems, or monitoring tools, the MCP + vault pattern is the production-safe approach. The post notes that many agent deployments fail because they lose access to a source without anyone noticing—proper credential and source management prevents this. Real-world example: an agent that provisions network connectivity needs access to IPAM, device APIs, and ticketing; MCP servers for each let the agent work without seeing passwords.

Read the original at claude.dev ↗

Cisco Dialog gives IT control over Webex agent actions

UC Today / Shashi Bellamkonda analysis · Oct 8, 2026 · Analysis

Cisco announced Dialog, a governance layer for agents in Webex; IT can control agent access, tool invocation, and approval workflows; MCP integration enables both Claude and OpenAI dot agents to operate under Cisco's access control policies.

Why it matters Dialog framework shows how infrastructure teams can govern multi-model agent deployments at enterprise scale—applicable pattern for federated NetOps agents across vendor tools and platforms.

The Dialog framework addresses the critical gap: enterprises want agent capabilities but need IT control over what agents can do, who can invoke them, and audit trails. Cisco's approach: agents (Claude, OpenAI dots, or customer-built) operate inside Webex, but all access to external systems flows through MCP integration points that Dialog controls. IT administrators define policies: which agents can call which tools, which users can invoke agents, sensitive operations require human approval. The framework supports layered autonomy—fully autonomous for low-risk actions, human-in-loop for higher-risk changes. For infrastructure ops, this is the governance model that makes agents acceptable in production: agents can diagnose and remediate, but policy gates prevent accidental cascading changes or unauthorized access. Dialog enters beta early 2027, but the architecture shows that IT governance isn't bolted on after—it's native to how agents execute. This matters for NetOps teams considering agentic platforms: governance should be part of the agent framework, not a separate compliance layer.

Read the original at shashi.co ↗

IETF adopts aggregate performance reporting format for email

Suped · Oct 9, 2026 · Primary source

The IETF MAILMAINT working group advanced draft-ietf-mailmaint-aprf-00 to working-group status on October 9, marking formal adoption of the Aggregate Performance Reporting Format specification. This Internet-Draft codifies DNS discovery, JSON reporting structures, and reputation categories for email security telemetry.

Why it matters Email operators and infrastructure teams can now influence standardization of performance metrics via the IETF process; early adoption feedback shapes final RFC before production deployment decisions.

Aggregate Performance Reporting (APRF) evolved from an individual submission (draft-brotman-aggregate-performance-reporting) into formal IETF MAILMAINT working-group document stream. The October 9 publication represents a critical gate: WG has accepted responsibility for the specification, enabling community review, wording refinement, and resolution of open questions before Standards Track advancement. Version 00 remains early working text; the draft expires April 12, 2027, and may be revised, replaced, or abandoned. The intended status is Standards Track. The specification introduces JSON structures for streaming email authentication and deliverability metrics (DKIM, SPF, DMARC aggregate reports, reputation signals) over DNS. This standardizes what vendors currently implement ad-hoc, reducing fragmentation and enabling mail operators to ingest telemetry from multiple sources into unified monitoring without custom parsing. Practical implication: teams should not yet deploy production DNS records or build fixed parsers around current syntax, as the format will evolve during WG review.

Read the original at suped.com ↗

Claude Haiku 5.5 tokenizer counts 30 percent more tokens

Anthropic / Claude Platform Docs / LLM Stats · Oct 7, 2026 · Primary source

Anthropic released Claude Haiku 5.5 on October 7 with 1M token context, configurable adaptive effort levels (low/medium/high/xhigh/max), and two-tier pricing: $0.10/$0.50 per million tokens for prompts ≤100K tokens, $0.50/$2.50 above. The new tokenizer counts ~30% more tokens for identical text versus Haiku 4.5, creating a hidden cost increase despite lower per-token pricing.

Why it matters For cost-sensitive, high-volume production workloads (classification, extraction, routing, subagents), you must re-benchmark cost-per-completed-task, not cost-per-token, due to tokenizer changes. Tier-based pricing at 100K tokens changes cost structure for long-context agentic tasks.

Claude Haiku 5.5 targets high-volume, latency-sensitive production tasks: classification, extraction, routing, live customer support, and coding subagents. It is the first Haiku-class model with adjustable reasoning effort, allowing operators to trade off response quality, latency, and token consumption in a single call. The effort parameter defaults to medium; low effort is tuned for simple, routine classification; higher settings (high/xhigh/max) justify ambiguous extraction and tool-planning steps. Context window expands 5x from Haiku 4.5 (200K→1M tokens); max output also doubles (64K→128K). However, the updated tokenizer inflates token counts by approximately 30% for the same input text, offsetting the per-token price reduction ($0.10 input vs. $1.00 for Haiku 4.5). Pricing bifurcates at 100K-token prompt length: short prompts stay cheap; prompts exceeding 100K tokens jump to $0.50 input/$2.50 output. Available on Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Knowledge cutoff: June 2026.

Read the original at platform.claude.com ↗
Nothing in today's briefing matches that.

Listening

Podcasts and talks
  • Packet Pushers

    HN845: Real Life Use Cases for BGP Monitoring Protocol (BMP)

    Bart Dorlandt discusses practical applications and architecture of BGP Monitoring Protocol for service providers, including data collection strategies, pre/post-policy filtering capture, and handling excessive prefix volumes in production networks.

  • ipSpace.net blog

    Please Respond: State of Network Automation Survey

    The Network Automation Forum is running its annual State of Network Automation Survey. Ivan Pepelnjak encourages practitioners involved in network automation to complete the survey to provide data on adoption patterns and operational realities.

  • ipSpace.net blog

    Netmiko vs Ansible: Does It Matter?

    Ivan Pepelnjak analyzes the choice between Netmiko (lightweight CLI library) and Ansible (orchestration framework) for device configuration, noting netlab now supports both approaches and will default to Netmiko for several platforms.

Vendor Radar

Last 7 days · arrows compare with the 7 before

What changed this week

Last 7 days vs the 7 before

Biggest moves

Trending topics

agent · MCP · SRE · automation · observability · Agentic AI · RAG · OpenTelemetry · AIOps · LLM · SASE · Kubernetes

Get the daily briefing

Today's 3 things that matter and every story with why it matters, in your inbox each morning. Free, and you can unsubscribe at any time. Prefer a reader? Follow the RSS feed.