An on-premises, GPU-accelerated virtual try-on, AI styling, and live-inventory solution from Metrum AI, deployed on the Supermicro SuperWorkstation A+ Server AS -2115HV-TNRT with four AMD Radeon™ AI PRO R9700S GPUs.


Prepared in collaboration with AMD and Supermicro.
Executive Summary
Apparel commerce runs on confidence, and most shoppers never get it. A garment that looks right on a model can fit and drape differently on a real body, and online there is no fitting room to close the gap. The result is hesitation, abandoned carts, and a return rate that quietly erodes apparel margin. The virtual try-on tools meant to solve this usually send a customer's photos and video to the cloud, which trades one problem for a harder one: privacy, consent, compliance, and data-ownership exposure.
Metrum AI built the Virtual Try-On solution to deliver a photorealistic fitting room entirely inside the retailer's environment. A shopper selects a profile and garments from the catalog. The solution extracts pose-diverse keyframes from a short profile video and uses local diffusion models to render the selected garments onto the shopper, assembling a photorealistic try-on slideshow with AI styling tips and generated music. A voice-enabled AI stylist answers questions, recommends outfits, and surfaces live inventory and pricing, so browsing turns into a confident purchase.
Everything runs on a Supermicro SuperWorkstation A+ Server AS -2115HV-TNRT fitted with four AMD Radeon™ AI PRO R9700S GPUs and the AMD ROCm™ software stack. Customer video and imagery stay within infrastructure the retailer controls and never reach a third-party cloud. Retailers gain three business advantages from that design: higher purchase confidence and fewer returns, a differentiated shopping experience that sets the storefront apart, and predictable cost with no per-inference cloud billing.
This brief describes the business problem, the solution, its architecture, and the outcomes retailers can expect. It is written for IT directors, CTOs, infrastructure architects, and line-of-business merchandising, e-commerce, and store-operations leaders evaluating on-premises AI for the shopping experience. The figures in this brief come from a Metrum AI demonstration on the reference configuration in Table 3, running the sample garments and videos bundled with the solution repository.
The Business Challenge
Apparel retail lives and dies on fit and confidence. When a shopper cannot tell how a garment will look before it arrives, the retailer absorbs the cost, through abandoned carts, through returns that must be inspected and restocked, and through the slow erosion of trust when the item looks different than expected. Most retailers still discover that mismatch only after the sale, when the package comes back.
Where Retailers Lose Money Today
Fit and styling gaps recur across retailers of every size, and each one carries a cost.
| Operational Gap | Business Consequence |
|---|---|
| Shoppers cannot visualize fit before buying | Uncertainty drives hesitation, abandoned carts, and "looks different than expected" returns that erode apparel margin. |
| Try-on depends on sending customer imagery to the cloud | Customer photos and video leave the retailer's control, creating privacy, consent, and data-ownership exposure that limits where the experience can be deployed. |
| Styling guidance is generic and unstaffed | Shoppers get no personalized help, so cross-sell, upsell, and confident purchases are left on the table. |
| Inventory and pricing are disconnected from the try-on moment | Shoppers cannot tell what is in stock, where, or at what price, so ready-to-buy intent stalls at the decisive moment. |
Table 1. Recurring gaps in the apparel shopping experience and their business consequences
Why Cloud-First AI Falls Short for Customer Imagery
Diffusion-based try-on is the natural answer to these gaps, and many vendors deliver it from the cloud. That architecture introduces three problems that grow with catalog and customer volume.
- Bandwidth and latency. Continuous upload of customer video and imagery consumes bandwidth and adds latency to an experience that has to feel instant on the shop floor.
- Variable cost. Per-inference and per-image pricing turns an experience feature into an operating expense that scales with every session.
- Loss of custody over customer likeness. Customer-identifiable imagery leaves the retailer's control, which raises privacy, consent, and data-residency questions in many jurisdictions.
For an experience meant to build trust at the moment of purchase, a customer's likeness is the last thing a retailer should be sending off site.
The Market Shift: Experience Replaces Guesswork
Three developments have made local, real-time virtual try-on practical rather than aspirational.
Workstation-class GPUs now carry production diffusion. Four AMD Radeon™ AI PRO R9700S GPUs put 128 GB of aggregate video memory in a single 2U node, enough to run distributed diffusion try-on, multi-model perception, and local language model serving at once rather than in shifts.
Open-weight models close the capability gap. Compact, efficient diffusion, embedding, and language models run locally, so photorealistic try-on and stylist reasoning no longer depend on a frontier model behind an API.
Agentic AI turns catalogs into conversations. Software agents read preference, catalog, and inventory data and deliver the styling and logistics guidance a human associate would, by voice and in real time.
Virtual try-on has been a cloud novelty. These shifts make it an experience the retailer owns, running on hardware in the store, with customer data that never leaves it.
Solution Overview
Virtual Try-On is a fully on-premises solution that turns a short customer video and a garment catalog into a photorealistic fitting-room experience. It extracts pose-diverse keyframes from the customer's profile video, renders the selected garments onto the shopper with local diffusion, assembles a styled slideshow with AI tips and generated music, and wraps the whole experience in a voice-enabled AI stylist that also answers live inventory and pricing questions.
A single voice-first dashboard serves shopper and associate alike. The Style Desk runs a natural conversation with the AI stylist and its recommendations; the Virtual Try-On View plays the original customer video beside the AI-generated try-on and slideshow; and live GPU telemetry shows the work happening on premises, in real time, on hardware the retailer owns.

Figure 1. The voice-first Virtual Try-On dashboard: AI stylist conversation, original video beside the generated try-on, live styling recommendations, and GPU telemetry
Customer profiles, garments, prices, and inventory shown are AI-generated and illustrative.
Capability Highlights
| Capability | What It Delivers |
|---|---|
| Diffusion-powered virtual try-on | Dual FASHN inference servers render photo-realistic garment composites from a short customer video, while DWPose selects the sharpest front, side, and back keyframes. Runs on four AMD Radeon™ AI PRO R9700S GPUs with the AMD ROCm™ stack. |
| AI-styled slideshow with dynamic music | AI-curated keyframes are stitched into a smooth crossfade presentation, with LLM styling tips and ACE-Step-generated music for a polished, on-brand shopping experience. |
| Voice-first agentic shopping | OpenClaw agents running on a local Qwen 3.6 27B model conduct natural voice-based preference discovery; Kokoro TTS delivers conversational recommendations, logistics updates, and feedback collection. |
| Preference-scoped catalog and visual search | Recommendations are scoped by the department and fit preferences the shopper selects, never by attributes inferred from imagery. BAAI/bge embeddings, DINOv3 visual embeddings, and a local Milvus index enable text and image product discovery without a query leaving the store. |
| Live inventory and pricing in context | A logistics agent answers stock, price, and fulfillment questions by branch at the moment of try-on, turning shopper intent into a confident purchase. |
| Session history and comparison | Complete try-on history is preserved through the session, with side-by-side comparison of multiple generations for confident purchase decisions. |
| On-premises model serving | Diffusion, embedding, speech, music, and language models run locally through vLLM and Lemonade Server on AMD ROCm™, so no customer data or inference leaves the retailer's network. |
Table 2. Capability highlights and the operational value each one delivers
How the Solution Works
The solution follows a four-stage pipeline that runs entirely on the local system.
Capture. The shopper selects a customer profile and the system ingests a short profile video. FFmpeg extracts frames and uniform sampling produces a set of candidate keyframes.
Select. DWPose scores the candidates and selects pose-diverse keyframes, front, side, and back, and a human-parsing model produces a body-segmentation mask for accurate garment placement. Both models run on the AMD Radeon™ AI PRO R9700S GPUs.
Render. Dual FASHN diffusion inference servers composite the selected garments onto the chosen keyframes, producing photo-realistic try-on images.
Assemble and converse. A slideshow assembler stitches the try-on frames into a crossfade video with LLM styling tips and ACE-Step music. In parallel, the agent layer runs the voice conversation, styling recommendations, live inventory and pricing, and feedback capture.

Figure 2. Solution workflow, from customer profiles and video through the agentic styling team and the diffusion try-on pipeline to the dashboard
When a shopper asks for something specific, the response arrives across the whole dashboard at once. The stylist agent narrates a recommendation aloud, the selected garments render onto the shopper in the try-on view, the logistics agent reports stock and price by branch, and the whole session is preserved so the shopper can compare generations side by side before deciding.
Solution Architecture
The architecture layers a video and pose-perception pipeline, a distributed diffusion try-on engine, an agentic reasoning layer, a vector-search catalog, and a voice-enabled web dashboard on top of AMD ROCm™ and a containerized service stack. Every layer runs on the same local system, and the deployment uses standard container tooling so operations teams manage it with tools they already know.

Figure 3. Solution architecture, from hardware and AMD ROCm through perception, diffusion, agents, catalog, and dashboard
Reference Configuration
The following configuration is the recommended starting point for this workload and the system Metrum AI used for the demonstration described in this brief.
| Component | Specification |
|---|---|
| System platform | Supermicro SuperWorkstation A+ Server AS -2115HV-TNRT, 2U single-processor rackmount workstation with PCIe 5.0 |
| GPUs | 4 x AMD Radeon™ AI PRO R9700S, 32 GB per GPU |
| Processor | AMD Ryzen™ Threadripper™ PRO 9995WX, 96 cores |
| System memory | 256 GB DDR5 |
| Storage | 2 TB NVMe SSD |
| Operating system | Ubuntu 24.04.4 with the 6.17 HWE kernel |
| GPU software stack | AMD ROCm™ 7.2.3 |
| Container runtime | Docker Engine 25.0 or later with Docker Compose v2.20 or later |
Table 3. Reference hardware and software configuration used for solution deployment and validation
The R9700S is the passively cooled variant of the Radeon AI PRO R9700, built for the directed chassis airflow of a rackmount system rather than the open airflow of a desktop workstation. That is why four of them fit a 2U node.
Business Outcomes
Retailers evaluate experience technology on conversion, returns avoided, and cost predictability. The solution addresses all three.
- A preview before the purchase. Shoppers see a photorealistic rendering of fit and style on their own body before buying, at the moment they are deciding.
- A direct answer to the returns driver. "Looks different than expected" is the return reason this addresses: the shopper sees the garment rendered on themselves rather than on a model.
- Full data sovereignty. Customer video and imagery stay within infrastructure the retailer controls and never reach a third-party cloud, which simplifies privacy and consent review and supports data-residency requirements.
- A differentiated shopping experience. A voice-first AI stylist paired with photorealistic try-on sets the retailer's storefront apart, online and in store.
- Personalization at scale. Preference-scoped recommendations and visual search give every shopper an attentive associate without adding headcount.
- Predictable economics. Inference runs on hardware the retailer owns, so cost does not scale with session volume or per-image cloud pricing.
Deployment Model and Next Steps
The solution deploys as a containerized stack on a single system, and a guided setup routine validates the GPU environment, generates configuration, builds the container images, launches every service, and seeds the garment catalog. Retailers typically move through three phases.
- Validate. Stand up the reference configuration and run the solution against the bundled sample garments and videos to confirm try-on quality, styling behavior, and the voice experience for your merchandising.
- Integrate. Connect your own catalog, garment imagery, and live inventory and pricing feeds, then tune keyframe count, diffusion quality, and stylist prompts to your brand and assortment.
- Scale. Extend the experience by adding nodes, one per storefront, each running the same containerized stack, with in-store kiosks served from the local node, and add agents for the workflows your operation values most, such as loyalty, cross-sell, or clienteling.
To scope a deployment for your storefront, or to arrange a technical walkthrough of the Virtual Try-On dashboard, contact Metrum AI or your AMD and Supermicro account teams.
Appendix A: Agent Layer Detail
The agent layer runs a coordinated team that turns a shopper's spoken preferences into styling recommendations, live logistics answers, and captured feedback. Each agent reads a scoped set of data and returns a structured result, and the results appear inline in the dashboard with no manual step required.
| Agent | Role |
|---|---|
| Try-On Orchestrator | Coordinates the agent team, delegating styling, logistics, and feedback tasks and sequencing the try-on generation. |
| Fashion Stylist Agent | Runs voice-based preference discovery and generates catalog styling recommendations scoped to the selected customer profile. |
| Logistics Agent | Enriches recommendations with live inventory, pricing, and fulfillment detail by branch. |
| Customer Feedback Agent | Captures shopper outfit and experience feedback through an inline star rating to improve future recommendations. |
Table 4. Agent roles in the on-premises reasoning layer
References
| 1. | AMD. "AMD Radeon AI PRO R9700 Graphics." https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9700.html |
| 2. | AMD. "AMD ROCm Documentation." https://rocm.docs.amd.com |
| 3. | Supermicro. "SuperWorkstation A+ Server AS -2115HV-TNRT." https://www.supermicro.com/en/products/system/superworkstation/2u/as-2115hv-tnrt |
| 4. | Metrum AI. Virtual Try-On solution repository documentation and release notes. https://github.com/metrum-ai/amd-radeon-virtual-tryon |
| 5. | Image sources: Virtual Try-On dashboard, workflow, and architecture visuals from the Virtual Try-On solution repository. Product imagery courtesy of AMD and Supermicro. |
Disclaimers
Performance
Performance varies by hardware and software configuration, including testing conditions, system settings, application complexity, data quantity, batch sizes, software versions, and libraries used. Any performance figures referenced in this document are provided for informational purposes only and should not be interpreted as a guarantee of actual performance.
Model Performance Limitations
This solution is a technology demonstration validated against bundled sample garments and customer videos. Diffusion-based try-on and pose detection have known accuracy boundaries. Try-on accuracy degrades with occluded poses, fast motion, low-contrast footage, or unusual garments. For other customer videos or garment imagery, accuracy is not guaranteed and may produce misalignment or visual artifacts. Results should be treated as indicative rather than authoritative.
Data and Legal
Garment images, catalog metadata, prices, inventory, customer profiles, and demonstration videos used in this solution are AI-generated, simulated, and illustrative. They do not represent real people, products, prices, or inventory and should not be relied upon for operational, legal, or real-world decision-making. Deployments that capture customer imagery may be subject to privacy, biometric, consent, and data-protection requirements that vary by jurisdiction. Operators should confirm compliance obligations before production use. All materials are supplied as-is without warranties of any kind.
AMD, the AMD Arrow logo, Radeon, Ryzen, Threadripper, ROCm, and combinations thereof are trademarks of Advanced Micro Devices, Inc. Supermicro, SuperWorkstation, and A+ Server are trademarks or registered trademarks of Super Micro Computer, Inc. All other product names are used for identification purposes only and may be trademarks of their respective owners.
METRUM AI INC. 2026 © All rights reserved.