An on-premises, GPU-accelerated work-instruction authoring solution from Metrum AI that harmonizes multi-vendor OEM manuals and plant SOPs into safety-gated, illustrated procedures, deployed on the Supermicro SuperWorkstation A+ Server AS -2115HV-TNRT with four AMD Radeon™ AI PRO R9700S GPUs.


Prepared in collaboration with AMD and Supermicro.
Executive Summary
Manufacturers lose time, quality, and safety margin to work instructions that cannot keep up. OEM documentation is fragmented across vendors, formats, and languages; static SOPs do not reflect the live state of the machine an operator is standing in front of; and the hard-won knowledge of senior engineers lives in their heads. The AI authoring tools meant to close these gaps usually require sending sensitive manuals and process data off the plant floor, trading a documentation problem for a security one.
Metrum AI built the Work Instruction Generator to author accurate, illustrated, safety-gated work instructions entirely inside the plant. Coordinated AI agents ingest multi-vendor OEM manuals and plant SOPs, normalize terminology to a single plant standard, author step-by-step procedures grounded in live machine state, and re-render vendor diagrams into one consistent illustration style. A separate rules-based interlock service, independent of the language models, evaluates LOTO, PPE, and fixture conditions and marks each step Pass or Blocked before it is rendered. Authoring is AI-assisted; the gate is not. Each step traces back to the source document it came from.
Everything runs on a Supermicro SuperWorkstation A+ Server AS -2115HV-TNRT fitted with four AMD Radeon™ AI PRO R9700S GPUs and the AMD ROCm™ software stack. No manuals, SOPs, or process data leave the plant network. Manufacturers gain three business advantages from that design: instructions that reflect the machine's current condition rather than a static document; authoring that collapses manual SOP reconciliation into a generated draft an engineer reviews; and preserved tribal knowledge at a predictable, cloud-free cost.
This brief describes the business problem, the solution, its architecture, and the outcomes manufacturers can expect. It is written for IT directors, CTOs, infrastructure architects, and line-of-business manufacturing, operations, industrial-engineering, and EHS leaders evaluating on-premises AI for the factory floor. The figures in this brief come from a Metrum AI demonstration on the reference configuration in Table 3, running the sample OEM manuals, SOPs, and EV battery-pack assembly procedures bundled with the solution repository.
The Business Challenge
Manufacturing runs on precise, current instructions. When operators work from outdated, mistranslated, or mismatched documentation, the plant pays for it, in scrap and rework, in safety incidents and near-misses, and in the weeks it takes to onboard a new operator. Most plants discover the gap only after a defect ships or an incident is logged.
Where Manufacturers Lose Money Today
Documentation and knowledge gaps recur across plants of every size, and each one carries a cost.
| Operational Gap | Business Consequence |
|---|---|
| OEM documentation is fragmented across vendors, formats, and languages | Engineers spend weeks reconciling manuals by hand, and operators work from inconsistent, sometimes mistranslated instructions. |
| Static SOPs do not reflect live machine state | Instructions ignore current interlocks, alarms, and readiness, so a step can be unsafe or invalid for the machine's actual condition. |
| Operator tribal knowledge is undocumented | Senior-engineer workarounds and fixes live in people's heads and are lost to turnover, retirement, and shift changes. |
| AI authoring tools send manuals and process data to the cloud | Proprietary process IP and safety documentation leave the plant network, creating privacy, security, and compliance exposure. |
Table 1. Recurring gaps in manufacturing work instruction and their business consequences
Why Cloud-First AI Falls Short for Plant Documentation
AI authoring is the natural answer to these gaps, and many vendors deliver it from the cloud. That architecture introduces three problems for a manufacturing plant.
- Process IP leaves the plant. Uploading proprietary OEM manuals, SOPs, and live process data exposes the plant's most sensitive intellectual property and complicates compliance.
- Cost scales with revisions. Per-inference and per-document pricing turns authoring into an operating expense that grows with every procedure and every revision.
- A remote service cannot see the machine. An instruction grounded in live station state depends on telemetry that stays inside the plant network.
For instructions that must be correct, current, and safe on the floor, the plant's process IP and its control-system data belong on the plant's own hardware.
The Market Shift: Live Authoring Replaces Static SOPs
Three developments have made local, state-aware work-instruction authoring practical rather than aspirational.
- Workstation-class GPUs now carry production multi-model inference. Four AMD Radeon™ AI PRO R9700S GPUs put 128 GB of aggregate video memory in a single 2U node, enough to run multilingual authoring, vision-language document understanding, and diffusion illustration at once rather than in shifts.
- Open-weight models close the capability gap. Compact, permissively licensed language, vision, and image models run locally, so translation, authoring, and technical illustration no longer depend on a frontier model behind an API. Every model in this solution ships under Apache 2.0, which leaves the deployment terms with the plant.
- Agentic AI turns documents into drafts. Software agents ingest, normalize, author, illustrate, and export, so an engineer reviews a generated procedure instead of reconciling manuals by hand.
Work instruction has been a static document that ages from the day it is written. These shifts make it a live artifact the plant owns, generated against current machine state, with process IP that never leaves the network.
Solution Overview
The Work Instruction Generator is a fully on-premises solution that turns fragmented OEM manuals, plant SOPs, and live machine telemetry into accurate, illustrated, safety-gated work instructions on demand. A multi-agent harmonization engine ingests multi-vendor documentation, normalizes terminology to a plant standard, authors steps against live station state, re-renders vendor diagrams into one consistent style, and exports traceable PDF and HTML documents.
A single dashboard serves operator and engineer alike. The Machine Floor streams live telemetry from each station with status gauges; a procedure selector scopes the workspace to a procedure, its station, and its source documents; the Step Viewer presents authored steps as cards with tools, parts, interlock and citation badges and a generated illustration per step; and a metrics panel shows live GPU and CPU utilization and throughput, on hardware the plant owns.

Figure 1. The Work Instruction Generator dashboard: live machine floor, authored steps with interlock and citation badges, per-step illustration, and GPU telemetry
Simulated data. Stations, vendors, and procedures shown are fictitious.
Capability Highlights
| Capability | What It Delivers |
|---|---|
| Multi-agent harmonization engine | Coordinated agents ingest multi-vendor, multi-language OEM manuals and plant SOPs and normalize terminology to a plant standard, translated and authored by local language models on four AMD Radeon™ AI PRO R9700S GPUs with the AMD ROCm™ stack. |
| AI-rendered technical illustrations | Every authored step gets its own FLUX.2 [klein] 4B illustration, giving one consistent visual style across all vendors and procedures. |
| Deterministic safety gating on live machine state | A rules-based interlock service, independent of the language models, evaluates station readiness against LOTO, PPE, and fixture conditions received over MQTT / Sparkplug B. It marks each step Pass or Blocked before the step is rendered. Generated content is never permitted to release a blocked step. |
| Tribal knowledge capture | Senior-engineer corrections and workarounds are captured through an in-place review chat, which regenerates the step text and illustration and preserves the knowledge for continuity. |
| Traceable export and document history | Every document exports to PDF or HTML with each step tracing back to its source, and generation history and reference documents stay one click away for audit or reuse. |
| On-premises model serving | Language, vision, and image models run locally through optimized AMD ROCm inference runtimes, so manuals, SOPs, and process data never leave the plant network. |
Table 2. Capability highlights and the operational value each one delivers
How the Solution Works
The solution follows a four-stage pipeline that runs entirely on the local system.
- Ingest. The ingestion pipeline extracts OEM manuals, images, warnings, and citations and translates multi-vendor, multi-language documentation into a plant standard, indexing it into a vector store for retrieval.
- Ground. The orchestrator resolves the selected procedure and step position and reads live machine state, including alarms, readiness, and the last completed step, from the station simulators. Retrieval applies plant terminology and SOP context to the authoring prompt.
- Author and gate. The instruction author agent creates state-aware operator steps. Separately, the rules-based interlock service evaluates readiness flags against LOTO, PPE, and fixture conditions and marks each step Pass or Blocked per station. The gate runs outside the agent pipeline and its decision cannot be overridden by generated content.
- Illustrate, review, and export. The technical illustration agent renders unified line art for each step, the review-and-feedback agent injects captured knowledge and the approval trail, and the compositor-and-export agent compiles HTML and PDF output and stores the artifacts.

Figure 2. Solution workflow, from OEM manuals, procedure selection, and machine state through the generation agents and the independent interlock gate to the dashboard
When a station enters a fault or an interlock is unmet, the response is visible across the dashboard. The affected step is marked Blocked with the failing prerequisite, the machine card raises a critical alarm, and the document withholds the step until the condition clears, so an operator never receives an instruction the machine's state cannot safely support.
Solution Architecture
The architecture layers a document-ingestion and retrieval pipeline, a multi-agent authoring and illustration engine, an independent rules-based interlock service, and a web dashboard on top of AMD ROCm™ and a containerized service stack. Live station telemetry arrives over an MQTT / Sparkplug B message bus. Every layer runs on the same local system, and the deployment uses standard container tooling so operations teams manage it with tools they already know.

Figure 3. Solution architecture, from hardware and AMD ROCm through ingestion, the OpenClaw agent control plane, models, and dashboard
Reference Configuration
The following configuration is the recommended starting point for this workload and the system Metrum AI used for the demonstration described in this brief.
| Component | Specification |
|---|---|
| System platform | Supermicro SuperWorkstation A+ Server AS -2115HV-TNRT, 2U single-processor rackmount workstation with PCIe 5.0 |
| GPUs | 4 x AMD Radeon™ AI PRO R9700S, 32 GB per GPU |
| AI model allocation | Qwen3.5-9B (GPU 0); Gemma-4-E2B-it x2 (GPU 1); FLUX.2 [klein] 4B, community GGUF quantization, x2 (GPU 2 and GPU 3). All models Apache 2.0 |
| Processor | AMD Ryzen™ Threadripper™ PRO 9995WX, 96 cores |
| System memory | 256 GB DDR5 |
| Storage | 2 TB NVMe SSD |
| Operating system | Ubuntu 24.04.4 with the 6.17 HWE kernel |
| GPU software stack | AMD ROCm 7.2.3 |
| Inference and serving | vLLM and Lemonade Server |
| Container runtime | Docker Engine 25.0 or later with Docker Compose v2.20 or later |
Table 3. Reference hardware and software configuration used for solution deployment and validation
The R9700S is the passively cooled variant of the Radeon AI PRO R9700, built for the directed chassis airflow of a rackmount system rather than the open airflow of a desktop workstation. That is why four of them fit a 2U node.
Business Outcomes
Manufacturers evaluate authoring technology on quality, labor recovered, and cost predictability. The solution addresses all three.
- Steps that match the machine's current condition. An operator receives a step only when the interlock service confirms the station's readiness conditions are met, rather than working from a static document written months earlier.
- Authoring effort moves from reconciliation to review. Manual SOP reconciliation across vendors and languages becomes a generated draft an engineer reviews, and new operators ramp on clear, illustrated, current instructions.
- Preserved tribal knowledge. Senior-engineer corrections are captured in place and carried forward, protecting operational know-how against turnover and retirement.
- Consistent, compliant documentation. One illustration style and one plant-standard terminology across every vendor, with each step traceable to its source for audit.
- Full data sovereignty. Proprietary manuals, SOPs, and process data never leave the plant network, which simplifies security and compliance review.
- Predictable economics. Inference runs on hardware the plant owns, so cost does not scale with procedure count, revisions, or per-inference cloud pricing.
Deployment Model and Next Steps
The solution deploys as a containerized stack on a single system, and a guided setup routine validates the GPU environment, generates configuration, builds the container images, launches every service, and ingests the OEM and SOP documents. Manufacturers typically move through three phases.
- Validate. Stand up the reference configuration and run the solution against the bundled sample manuals, SOPs, and EV battery-pack assembly procedures to confirm authoring quality, interlock behavior, and illustration style.
- Integrate. Connect your own OEM manuals, plant SOPs, and live machine telemetry over MQTT / Sparkplug B, then tune terminology, interlock rules, and quality gates to your plant standard.
- Scale. Extend across stations and lines, and add agents for the workflows your operation values most, such as change management, audit packaging, or multilingual operator delivery.
To scope a deployment for your plant, or to arrange a technical walkthrough of the Work Instruction Generator dashboard, contact Metrum AI or your AMD and Supermicro account teams.
Appendix A: Agent Layer Detail
The agent layer runs a coordinated team that turns fragmented documentation and live machine state into a safety-gated, illustrated, traceable document. Each agent reads a scoped set of data and returns a structured result, and the results are assembled into the final work instruction with no manual reconciliation required. The interlock service is not part of the agent team. It runs as a separate rules-based service, evaluates readiness flags against LOTO, PPE, and fixture conditions, and marks each step Pass or Blocked. The agents below author and assemble the document; they do not evaluate safety conditions.
| Agent | Role |
|---|---|
| Work Instruction Orchestrator | Routes the state-aware pipeline and sequences the agent team. It does not evaluate safety conditions. |
| Document Ingestion and Translation | Extracts OEM manuals, images, warnings, and citations and translates multi-vendor, multi-language documentation to the plant standard. |
| SOP Reference Agent | Retrieves SOP, terminology, and quality guidance from the indexed plant standard to ground the authored document. |
| Instruction Author Agent | Creates state-aware, multilingual operator steps grounded in machine state and source documents. |
| Technical Illustration Agent | Renders unified line art for each step from photos, blueprints, and schematics. |
| Review and Feedback Agent | Injects captured tribal knowledge and the approval trail into the document. |
| Compositor and Export Agent | Compiles HTML and PDF output and stores the traceable artifacts. |
Table 4. Agent roles in the on-premises authoring layer
References
| 1. | AMD. "AMD Radeon AI PRO R9700 Graphics." https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9700.html |
| 2. | AMD. "AMD ROCm Documentation." https://rocm.docs.amd.com |
| 3. | AMD. "Lemonade Server - local LLM serving." https://github.com/lemonade-sdk/lemonade |
| 4. | Supermicro. "SuperWorkstation A+ Server AS -2115HV-TNRT." https://www.supermicro.com/en/products/system/superworkstation/2u/as-2115hv-tnrt |
| 5. | Metrum AI. "Work Instruction Generator solution repository." https://github.com/metrum-ai/amd-radeon-work-instruction |
| 6. | Alibaba. "Qwen large language models." https://github.com/QwenLM/Qwen3 |
| 7. | Google. "Gemma open models." https://ai.google.dev/gemma |
| 8. | Black Forest Labs. "FLUX image generation models." https://bfl.ai/models |
| 9. | BAAI. "BGE / FlagEmbedding - text embeddings." https://github.com/FlagOpen/FlagEmbedding |
| 10. | Zilliz. "Milvus - vector database." https://milvus.io |
| 11. | vLLM. "High-throughput LLM inference and serving." https://docs.vllm.ai |
| 12. | NATS. "Connective technology / messaging system." https://nats.io |
| 13. | Eclipse Foundation. "Mosquitto - MQTT broker." https://mosquitto.org |
| 14. | Eclipse Foundation. "Sparkplug specification." https://sparkplug.eclipse.org |
| 15. | Prometheus. "Monitoring and time-series metrics." https://prometheus.io |
| 16. | Image sources: Work Instruction Generator dashboard, workflow, and architecture visuals from the solution repository. Product imagery courtesy of AMD and Supermicro. |
Disclaimers
Performance
Performance varies by hardware and software configuration, including testing conditions, system settings, application complexity, data quantity, batch sizes, software versions, and libraries used. Any performance figures referenced in this document are provided for informational purposes only and should not be interpreted as a guarantee of actual performance.
Model Performance Limitations
This solution is a technology demonstration validated against bundled sample OEM manuals, SOPs, and EV battery-pack assembly procedures. Language, vision, and diffusion models have known accuracy boundaries; generated work instructions, translations, and illustrations may contain errors and must be independently reviewed. Generated work instructions have not been validated against a real safety program and must not be used on an actual production floor without independent engineering and safety review. The interlock service in this demonstration is not a safety-rated control function and has not been assessed against ISO 13849 or IEC 62061. It supplements, and does not replace, the machine's own safety system.
Data and Legal
The OEM manuals, SOPs, station telemetry, vendor names, prices, and assembly procedures used in this solution are simulated and illustrative. They do not represent a real manufacturer, plant, or production line and should not be relied upon for operational, legal, or real-world decision-making. Deployments that process proprietary manufacturing documentation and control-system data may be subject to safety, employment, export-control, and data-protection requirements that vary by jurisdiction. Operators should confirm compliance obligations before production use. All materials are supplied as-is without warranties of any kind.
AMD, the AMD Arrow logo, Radeon, Ryzen, Threadripper, ROCm, and combinations thereof are trademarks of Advanced Micro Devices, Inc. Supermicro, SuperWorkstation, and A+ Server are trademarks or registered trademarks of Super Micro Computer, Inc. All other product names are used for identification purposes only and may be trademarks of their respective owners.
METRUM AI INC. 2026 © All rights reserved.