The GPU-native runtime
for production AI.
Production AI is expensive, ungoverned, and invisible. GPU fleets idle at 30%. Agents act without audit. Inference bills arrive with no cost attribution. Mandarum is the GPU-native runtime that fixes all three — compile, serve, route, govern, and observe AI on NVIDIA accelerated computing, from H100 and Blackwell datacenters to Jetson AGX Orin at the edge. One control plane. Every model, agent, and GPU.
Inference sprawl is the new shadow IT.
Enterprises run dozens of models and agents across scattered GPU clusters, clouds, and edge deployments — with no shared serving layer, no utilization governance, no cost attribution, and no audit trail. GPUs sit at 30–40% utilization while inference bills compound. Agents act without identity. And at the factory floor, every unaudited Jetson inference is a liability waiting to be discovered.
Every inference request is authenticated
Per-model and per-agent service identities, scoped API keys, and short-lived credentials — enforced at every call across cloud H100 clusters and Jetson edge nodes. No anonymous inference.
GPU spend caps, guardrails, agent-graph enforcement
A declarative policy engine that runs before each inference call — and across entire agent task sessions. Block unauthorized model access. Enforce GPU budget limits per call and per task. Escalate to humans before consequential actions.
Every GPU cycle traced, every dollar attributed
OpenTelemetry + DCGM GPU hardware telemetry — from H100 SXM5 NVLink throughput to Jetson AGX Orin power draw. Token throughput, GPU utilization, p99 first-token latency, and $/1M tokens attributed per model, per team, and per request.
Five products. One GPU Runtime.
Each product is designed to work alone. All five are designed to work as one. Every governed inference feeds an optimization flywheel. Every telemetry trace flows into a single control plane. Every policy you write once applies at the H100 rack and the Jetson node. Adopt Voltra today. Add Praxon when you ship agents. The rest of the runtime is there when you need it.
Mandarum Voltra
A NeMo-trained routing/optimization model that learns optimal GPU endpoint, quantization tier, and batching strategy from your fleet's own inference logs. Cuts $/token with every request served.
Explore Voltra →Mandarum Praxon
Agent-graph-aware governance — task-level budget caps, scope enforcement, and NeMo Guardrails across entire multi-step agent sessions, not just individual calls. Reasoning-chain audit trail for EU AI Act.
Explore Praxon →Mandarum Cuvex
Co-located GPU pipeline: NeMo Retriever embedding → RAPIDS cuVS billion-scale ANN with ACL-native enforcement → GPU reranking. No CPU round-trip. Sub-10ms retrieval at any corpus size.
Explore Cuvex →Mandarum Edgeon
Runtime + Sentinel on NVIDIA Jetson AGX Orin for manufacturing, robotics, and healthcare. Same identity, policy, and audit as your cloud H100 fleet — now governing every on-device inference action at the factory floor. IEC 62443, ISO 26262, and FDA SaMD compliance packs built in.
Explore Edgeon →Mandarum Twinex
Build a live digital twin of your GPU fleet from DCGM telemetry. Simulate new model deployments, hardware upgrades (H100 → H200), and Dynamo topology changes — before committing a dollar of compute.
Explore Twinex →One platform. Every model, agent, and GPU fleet.
No lock-in on models. No lock-in on agents. No lock-in on hardware. Mandarum wraps what you already run — NIM microservices, Triton-served open-weight models, LangGraph, CrewAI, or the OpenAI Agents SDK — while a NeMo-trained routing model continuously optimizes endpoint selection across your fleet. Bring your GPUs. Bring your models. Keep your data.
Triton → TensorRT-LLM → NIM → Dynamo → Voltra
Disaggregated prefill/decode via NVIDIA Dynamo. TRT-LLM compilation for 30–50% latency reduction per model. Voltra's NeMo-trained routing model selects the optimal GPU endpoint and quantization tier per request — improving with every inference served.
Jetson AGX Orin at the factory floor
Edgeon governs every on-device inference action on NVIDIA Jetson — manufacturing vision AI, industrial robotics, and medical imaging — with the same identity, policy, and immutable audit as your H100 cloud fleet.
Task-level policy, not just per-call
Praxon enforces budget caps and scope boundaries across entire agent task sessions — NeMo Guardrails + session-level context analysis + human-in-the-loop escalation before consequential actions.
GPU-native, ACL-enforced recall
Cuvex co-locates embedding, cuVS billion-scale ANN with ACL-native enforcement, and NeMo Retriever reranking on GPU — no CPU round-trip, sub-10ms retrieval at any corpus size.
Simulate before you deploy
Twinex builds a live DCGM-fed digital twin of your GPU fleet. Simulate new model deployments, H100→H200 upgrades, and Dynamo topology changes — before a dollar of compute is committed.
A real inference request, end to end.
Run a sample production request: TRT-LLM compiled model, Voltra routing to the cheapest GPU tier, Praxon session policy enforced, Cuvex RAG grounding, DCGM telemetry — all in one governed trace.
Production-grade from day one.
One pane. Every model, agent, and Jetson node.
Drill from a GPU utilization alert on an H100 cluster or a Jetson factory node to a specific inference request. Replay any agent session. Diff any model version. Export $/1M tokens per team to your data warehouse.
Open by design. GPU-native by architecture.
Bring any model
Anthropic, OpenAI, Google, open-weight via NIM/Triton/vLLM, or your NeMo fine-tuned models. Voltra routes each request to the optimal GPU tier and quantization level automatically.
Bring any agent framework
LangGraph, OpenAI Agents SDK, CrewAI, MCP servers. Praxon wraps any framework with task-level governance, NeMo Guardrails, and reasoning-chain audit — without rewriting your agent code.
Bring your full GPU fleet
NVIDIA H100/H200/Blackwell datacenter clusters, multi-cloud, on-prem, BYOC, and Jetson AGX Orin at the edge. One control plane, one audit trail, one $/1M-token cost view — cloud to factory floor.
Stop deploying models. Start governing them — cloud to edge.
The inference era is defined by efficiency and governance — not raw model capability. Join the teams deploying production AI on NVIDIA accelerated computing with Mandarum. White-glove onboarding, direct access to the founding team, and a 30-day pilot with documented GPU cost reduction.