
Policy-governed, multi-agent infrastructure operations on Dell PowerEdge™ R7725 server with AMD Instinct™ MI350P PCIe card, AMD EPYC™ processors, Dell PowerSwitch™ networking, and Metrum AI Infra Agents: monitoring, reasoning, and remediation that run on-premises.

Autonomous drone swarm inspecting live server racks at Dell Technologies World 2026 (concept rendering)
EXECUTIVE SUMMARY
The AI boom has set off a historic data center buildout, with capital spending projected to reach roughly $1.7 trillion a year by 2030. Modern AI infrastructure is denser, more distributed, and more operationally complex than anything that came before it, and that complexity is expensive: an estimated $400 billion a year is already lost to unplanned downtime across the Global 2000. This investment only returns value while the infrastructure is running. Delivering that ROI now requires 24x7 continuous monitoring and autonomous remediation that operate at the speed of the infrastructure itself.
Dell Technologies, AMD, and Metrum AI have built a locally deployed, policy-governed multi-agent platform for infrastructure operations. It runs on Dell PowerEdge with AMD Instinct acceleration and connects to Dell management and networking interfaces, so infrastructure agents can detect an incident, reason over telemetry, confirm ambiguous physical conditions, and initiate approved remediation without depending on cloud-hosted inference.
This is not a replacement for operations teams or Dell management tools. It is a Dell-based reference architecture for making those teams faster, more consistent, and better equipped to operate AI infrastructure at scale. Dell provides the enterprise infrastructure foundation, including PowerEdge compute, iDRAC management, and Dell PowerSwitch networking. AMD provides accelerated local compute for reasoning and vision. Metrum AI provides multi-agent intelligence that monitors, reasons, escalates, and remediates within enterprise guardrails. The proof of concept was demonstrated on live Dell hardware at Dell Technologies World 2026.
AT A GLANCE
What this solution delivers
- Enterprise AI infrastructure foundation from Dell. Dell PowerEdge R7725 with AMD Instinct MI350P PCIe GPUs and AMD EPYC 9965 processors, connected to Dell PowerSwitch S5248F-ON networking - a path to bring agentic AI operations into existing on-premises environments without redesigning the facility.
- Local agentic inference. Reasoning, vision analysis, and operator chat run on the server using AMD Instinct acceleration, reducing dependence on external cloud inference during infrastructure events.
- Dell management and networking integration. Agents act through iDRAC Redfish, SONiC-based network controls, and ITSM workflows, making Dell PowerEdge and PowerSwitch the operational foundation for governed remediation.
- Closed-loop autonomous recovery. A detect, diagnose, confirm, remediate, and recover workflow: a direct path where telemetry is sufficient, and physical inspection where it is ambiguous.
- Physical confirmation when telemetry cannot prove root cause: For ambiguous rack-level conditions, the platform dispatches an autonomous Crazyflie drone swarm as a native agent tool call. Imagery is analyzed on-chassis by the vision-language model and attached to the incident record - a thirty-second inspection in place of a walk-down.
- Demonstrated on live Dell hardware as a proof of concept. The Metrum AI Infra Agents platform ran on Dell infrastructure; drone inspection and simulated fault recovery were shown as proof-of-concept workflows, not benchmarks.
| Partner | Contribution | Customer Value |
| Dell Technologies | PowerEdge compute, Dell PowerSwitch S5248F-ON networking, iDRAC management, validated infrastructure design, lifecycle support, enterprise deployment experience. | The trusted on-premises foundation for running and operating agentic AI infrastructure at scale. |
| AMD | AMD Instinct MI350P PCIe GPU acceleration (144 GB HBM3E, up to 4 TB/s bandwidth) and AMD EPYC 9965 CPU performance (192 cores per processor). | Local reasoning, vision, retrieval, orchestration, and operator interaction on accelerated infrastructure. |
| Metrum AI | Multi-agent monitoring, reasoning, policy enforcement, remediation workflows, and operator-facing incident context. | The intelligence layer that detects, confirms, escalates, remediates, and records infrastructure events. |
THE PROBLEM
The gap between detection and resolution
AI infrastructure is now a business-critical production asset. Global data center capital spending is on track to reach roughly $1.7 trillion a year by 20301, yet the systems built at that scale still fail. Unplanned downtime costs the Global 2000 an estimated $400 billion annually2, with human error and slow human response the single largest cause.
Detection is rarely the hard part. Most environments already generate alerts from servers, GPUs, storage, networks, and applications through existing management tools, including Dell iDRAC and OpenManage. The hard part is converting those signals into fast, accurate, policy-compliant action when symptoms are ambiguous. A loose cable, a failing transceiver, a switch-port fault, a firmware drift, and a configuration change can all produce similar symptoms, and acting on the wrong root cause can deepen the outage.
That ambiguity creates an operational gap. Alerts enter a queue, engineers correlate telemetry, someone validates the cause, a remediation path is chosen, and evidence is added to the ticket. When the root cause is unclear, industry surveys cite resolution windows that routinely extend to hours4, and for 54% of operators a single significant incident exceeds $100,0003. Much of that cost accumulates while engineers work through ambiguous symptoms, with capital depreciating without producing return.
Metrum AI Infra Agents extend that operational workflow on the Dell platform. The agents reason over telemetry from existing Dell management and monitoring sources, determine whether physical confirmation is required, invoke inspection when needed, and execute or recommend remediation through governed workflows and the same Dell interfaces operators already use. In the DTW 2026 proof of concept, this detect-confirm-remediate-recover loop ran on the Dell reference infrastructure demonstrated at the event.
THE DELL PLATFORM
Dell PowerEdge + MI350P: Enterprise on-prem agentic AI
Autonomous infrastructure operations require more than an AI model. They require an enterprise platform that runs inference locally, connects to real infrastructure control planes, supports governed operations, and fits the data centers customers already operate. Dell PowerEdge provides that foundation.
The agentic operations stack runs on Dell PowerEdge R7725 with AMD Instinct MI350P PCIe GPUs and AMD EPYC processors, connected to Dell PowerSwitch networking and managed through familiar interfaces including iDRAC Redfish and SONiC-based network controls. Enterprise customers do not want autonomous operations to depend on one-off lab configurations. They need a repeatable pattern that moves from demonstration to pilot to production with clear operational ownership. Dell brings the platform discipline that transition requires: PowerEdge server design, management consistency through iDRAC across the fleet, PowerSwitch networking, lifecycle and firmware/BIOS practices, services, and support.
Keeping intelligence close to the infrastructure it operates reduces dependence on external cloud inference during an incident while supporting control, data locality, and resilience. That is Dell's strategic role: turning agentic AI from a software concept into a supportable infrastructure capability.
Why Dell PowerEdge with PCIe MI350P matters for enterprise deployment
AMD Instinct MI350P in PCIe form factor is purpose-built for organizations that need agentic and generative AI in existing air-cooled data centers. Installed directly in Dell PowerEdge R7725 servers, it delivers drop-in GPU density without facility redesign, with up to 4 TB/s memory bandwidth per published AMD specifications. It is the realistic starting point for enterprises moving agentic AI from pilot to production on-premises.
The Dell PowerEdge R7725 is engineered for exactly that. It is a 2U, dual-socket server built for dense AI compute and high-capacity NVMe storage. The PCIe-based AMD Instinct MI350P installs into standard, air-cooled PowerEdge servers as a Dell PowerEdge reference configuration demonstrated at Dell Technologies World 2026, without re-architecting the data center and with no dependency on cloud inference. The Dell Technologies World 2026 demonstration used this as a reference configuration on Dell PowerEdge hardware. An organization already running PowerEdge on-premises can extend it for agentic AI operations without re-architecting the data center.
The R7725 hosts both the intelligence that operates the environment and the infrastructure being operated, managed through one consistent Dell iDRAC Redfish interface. The same out-of-band control plane that ran this live demonstration is the same interface customers use across their PowerEdge fleet, making this a reference configuration, not a one-off lab setup.
In this reference configuration the agent runtime executes on a Dell PowerEdge R7725 acting as the control-plane node, reaching managed infrastructure through the iDRAC Redfish out-of-band path and the SONiC fabric. Because management runs out of band, the agents retain visibility and control even when the in-band data path or a managed node's operating system is degraded. Production deployments can place the agent runtime on a dedicated management node to isolate it from the failure domains it operates on.
The agents do not replace Dell iDRAC, OpenManage, or existing monitoring and ticketing systems. They consume alerts and telemetry from those systems and orchestrate the next steps: correlation, physical confirmation when telemetry is insufficient, and governed remediation through the same Dell interfaces operators already trust.


Three workloads, one MI350P on Dell PowerEdge
Autonomous infrastructure operations require three kinds of inference at once: language reasoning over telemetry and logs, vision-language analysis of inspection imagery, and operator chat against the same incident context. The MI350P's 144 GB of HBM3E is what makes running all three on a single card practical. In this configuration we deployed Qwen3.6-35B-A3B (35B-parameter MoE vision-language model, ~3B active parameters per token) via vLLM on AMD ROCm 7.2 at tensor parallel 1 as a unified model for reasoning, drone imagery interpretation, and operator chat, with all three workloads resident concurrently on one in-rack GPU. There is no model swapping and no cloud round-trip. The large memory footprint is the enabling factor. It lets the full multi-agent workload stay on one in-rack GPU rather than being split across nodes or offloaded to a remote endpoint.
We make no performance claims for this configuration. The point is architectural: what the memory footprint allowed us to deploy on-chassis in this proof of concept. This is not a measure of throughput, latency, or production performance.

AMD Instinct MI350P · CDNA 4 · 144 GB HBM3E
- LLM reasoning: correlates telemetry, logs, and switch events to classify the fault.
- VLM vision: reads drone imagery (unseated cable) in the same session.
- Operator chat: the operator chat interface answers questions against the same incident context.
| Attribute | MI350P (published spec) |
| Architecture | AMD CDNA 4; 128 compute units, 512 matrix cores |
| Memory | 144 GB HBM3E on a 4096-bit bus, up to 4 TB/s bandwidth; 128 MB Infinity Cache, full-chip ECC |
| Form factor | Full-height, full-length dual-slot PCIe 5.0 x16; passively cooled for air-cooled servers |
| Power | 600 W TBP* |
| Precision support | Native MXFP4, MXFP6, FP8; sparsity support for mainstream 8-bit and 16-bit precisions (per published AMD specifications) |
| Scale | Up to 8 cards per node (~1,152 GB aggregate HBM3E). The DTW 2026 reference configuration on PowerEdge R7725 used a single MI350P. |
*Published AMD specification. The Dell PowerEdge R7725 supports double-width GPUs at up to 450 W.
THE INTELLIGENCE
Metrum AI Infra Agents: Running on Dell PowerEdge
On top of that hardware run the Metrum AI Infra Agents: a locally deployed, cloud-independent AIOps platform for inference that delivers full-lifecycle automation across multivendor compute, network, storage, and VM environments, and extends to custom infrastructure and playbooks. On the DTW 2026 configuration, the platform runs entirely on the R7725's MI350P and operates the full incident lifecycle: detect, confirm, remediate, and recover, without escalating to a human.
Autonomous action is governed by the platform's built-in Know Your AI (KYAI) framework, which enforces policy guardrails, configurable human-in-the-loop approval gates, and a fully traceable audit trail for every action the agents take.
The Metrum AI Infra Agents is a deployable platform. At Dell Technologies World 2026, drone-based physical inspection and simulated fault injection were shown as proof-of-concept workflows on live hardware, not as production benchmarks.
Because every model runs locally on the server's MI350P, the agents do not depend on a network path to a cloud endpoint to think. The reasoning that classifies a fault, the vision that reads a drone's camera feed, and the chat that answers an operator's question all execute on the same server being repaired. That is precisely what you want when the failure may lie in the network itself.
And these are agents, not scripts. A monitoring agent continuously watches server health, GPU state, storage, and network behavior. When something looks wrong, a reasoning agent classifies the fault and decides whether telemetry is sufficient or visual confirmation is required. A specialist agent then executes the specific fix the confirmed root cause calls for, escalating through roles such as Operations Manager and Network Specialist as shown in the workflow examples below. The distinction matters: threshold automation applies the same predefined response to a symptom, while an agent selects a response appropriate to the cause it has actually confirmed.
Physical confirmation is what makes autonomous remediation safe rather than risky. When drone imagery shows a loose cable instead of a failed NIC, the agent reroutes traffic through SONiC failover rather than reloading a driver. That visual check sits between classification and action as a decision gate. Operators stay informed through OpenClaw, a chat interface that presents every incident as a structured thread; where a workflow needs sign-off, the KYAI approval gates can require human sign-off before any action runs.
PHYSICAL AI
The drone that sees what telemetry can't
Some faults cannot be confirmed from telemetry alone. Packet loss may indicate a link problem without distinguishing a loose cable from a failing transceiver, a switch-port issue, or a configuration error. Conventionally, that uncertainty sends a person to walk the aisle. In a continuously running facility, that walk-down is the bottleneck.
The platform invokes physical inspection through the same agent interface that reads telemetry and drives iDRAC and SONiC. When a fault is ambiguous, the inspection agent dispatches an autonomous Crazyflie drone swarm as a native tool call. The drones navigate by Lighthouse positioning: centimeter-accurate, no GPS, using local Lighthouse base stations for indoor positioning, with no need for ambient light. Qwen3.6-35B-A3B, running via vLLM on AMD ROCm 7.2 at tensor parallel 1 on the same Dell PowerEdge chassis, interprets the imagery, closing the loop on-premises.

The swarm inspects the rack face; the vision-language model interprets the imagery on the server (concept render).

The inspection loop: the agent dispatches, the swarm flies and films, the VLM reads it, the agent acts: one closed loop that runs entirely on the server.
This is best understood as an advanced confirmation mechanism, not the core value by itself. A fault that once required a person on the floor becomes a thirty-second flight, and inspection coverage scales with swarm size rather than headcount. If imagery confirms a loose cable rather than a failed NIC, the agent selects a safer path. It reroutes traffic while a field technician is dispatched, and the drone-confirmed imagery attaches to the ServiceNow ticket as an auditable record. A Unitree Go2 quadruped and G1 humanoid extend the same dispatch primitive for ground-level and visitor-facing roles. Production deployments would evaluate physical-inspection workflows against each facility's safety, security, and compliance requirements.
HOW IT FITS TOGETHER
One system, three layers
At a glance: operators and ITSM on top, the Metrum AI Infra Agents in the middle running locally on Dell PowerEdge with AMD Instinct MI350P, and the Dell / AMD infrastructure platform underneath. The drone swarm and physical-AI fleet are invoked through the same tool-call interface as iDRAC and SONiC.

Solution architecture - the agents run locally on the Dell platform they manage.
INCIDENT WORKFLOW EXAMPLES
Two faults, two different decisions
The same platform handles these completely differently, because it reasons about whether it can trust telemetry or needs to see the rack with its own eyes.
Demo status at Dell Technologies World 2026
| Element | Status |
| Dell infrastructure foundation | Demonstrated on Dell PowerEdge R7725 with AMD Instinct MI350P PCIe GPUs, AMD EPYC 9965 processors, and Dell PowerSwitch S5248F-ON networking. |
| Metrum AI Infra Agents | Demonstrated as the local multi-agent operations layer running on Dell infrastructure. |
| Drone-based physical inspection | Proof-of-concept tool call for ambiguous physical fault confirmation. |
| Simulated fault injection and recovery | Proof-of-concept workflow on live hardware. |
| Benchmark performance | Not claimed. Not an MTTR, availability, or throughput benchmark. |
| Guaranteed autonomous operation | Not claimed. Production uses customer-defined policy, approval gates, and audit controls. |
Scenario 1: Network fault (drone dispatched)
- Detect & Diagnose: Level 1 Support Agent flags packet loss on the active path.
- Confirm: L1 Agent + inspection drone; the VLM confirms a loose cable, not a NIC fault.
- Escalate: L1 to Operations Manager to Network Specialist, with full context.
- Resolve: The Network Specialist Agent initiates a SONiC-based failover to restore service and minimize disruption while preserving the incident record.
- Engage: A Field Specialist is dispatched to reseat the cable, arriving with the drone-confirmed imagery.
Scenario 2: CPU / BIOS drift (direct remediation)
- Detect & Diagnose: Level 1 Support Agent identifies CPU throttling after a patch reboot.
- Escalate: L1 to Operations Manager to Hardware Agent; telemetry confirms the cause, so no drone is needed.
- Resolve: Systems Admin Hardware Agent reapplies the performance BIOS profile via iDRAC.
References
| 1. | Dell'Oro Group, "AI Boom Drives Data Center Capex to $1.7 Trillion by 2030," 2025. [Online]. https://www.delloro.com/news/ai-boom-drives-data-center-capex-to-1-7-trillion-by-2030/ |
| 2. | Splunk and Oxford Economics, "The Hidden Costs of Downtime," 2024. [Online]. https://www.splunk.com/en_us/newsroom/press-releases/2024/conf24-splunk-report-shows-downtime-costs-global-2000-companies-400-billion-annually.html |
| 3. | Uptime Institute, "Annual Outage Analysis 2024," 2024. [Online]. https://uptimeinstitute.com/resources/research-and-reports/annual-outage-analysis-2024 |
| 4. | Traversal, "What is MTTR?" [Online]. https://www.traversal.com/blog/what-is-mttr |
| 5. | Uptime Institute, "2022 Global Data Center Survey Reveals Strong Industry Growth," 2022. [Online]. https://uptimeinstitute.com/about-ui/press-releases/2022-global-data-center-survey-reveals-strong-industry-growth |
Disclaimer
This document is provided for informational purposes only. It is a proof-of-concept technical brief, not a product announcement, white paper, or benchmark. Metrum AI developed this proof of concept in partnership with Dell Technologies, AMD, and Broadcom and demonstrated it on live Dell hardware at Dell Technologies World 2026. The autonomous drone swarm dispatch and simulated, injected faults described herein were shown to illustrate what becomes possible when multi-agent infrastructure operations run on Dell PowerEdge with AMD Instinct acceleration and Dell PowerSwitch networking. They are not guarantees of performance, availability, or outcomes in other environments. Actual results will vary based on workload, configuration, deployment conditions, and other factors.
Autonomous remediation requires customer-defined policy, approval gates, rollback procedures, credential controls, and change-management integration. Production deployments must be evaluated against each organization's safety, security, and compliance requirements.
Statements regarding future capabilities, roadmaps, or anticipated benefits are forward-looking and subject to change without notice. Nothing in this document constitutes a commitment, warranty, or guarantee of any kind, and no part of it should be relied upon in making a purchasing decision.
Dell, PowerEdge, PowerSwitch, and iDRAC are trademarks of Dell Inc. or its subsidiaries. AMD, EPYC, Instinct, CDNA, and ROCm are trademarks of Advanced Micro Devices, Inc. Broadcom and SONiC-related marks are the property of Broadcom Inc. ServiceNow is a trademark of ServiceNow, Inc. Crazyflie is a trademark of Bitcraze AB. Unitree, Go2, and G1 are trademarks of Unitree Robotics. All other trademarks are the property of their respective owners. Use of these names does not imply endorsement.
Third-party statistics and projections cited herein are attributed to their respective sources and have not been independently verified by Metrum AI.