Mandarum Platform Simulate · GPU Fleet Digital Twin
Mandarum Twinex logo

Mandarum Twinex — Model the fleet before you buy it.

A digital twin of your GPU inference fleet, built from real DCGM telemetry and rendered in NVIDIA Omniverse. Run what-if scenarios — hardware swaps, Dynamo disaggregation, batching, autoscaling, routing — and predict p99, throughput, utilization, and $/token before anything touches production.

NVIDIA Omniverse DCGM-calibrated What-if planning
Why Twinex exists

Why GPU decisions are still made on gut and guesswork.

01 · Guesswork

GPU buys are a gamble

Teams size fleets on spreadsheets and hope. Twinex predicts the impact of adding Blackwell or doubling traffic on p99 and $/token.

02 · Risky changes

Topology changes hit prod blind

Switching to disaggregated serving or new batching can backfire. Twinex tests the change in simulation first.

03 · Cost surprises

The GPU bill is a black box

Finance can't forecast inference spend. Twinex projects $/token across scenarios so capacity and budget decisions are grounded.

Capabilities

What Twinex does.

01 / Twin

A calibrated twin of your fleet.

Twinex ingests DCGM traces and serving logs to build a digital twin that replays GPU kernel timing, KV-cache, and batching behavior, calibrated to your real hardware and rendered in Omniverse.

  • DCGM + serving-log calibration
  • GPU-accurate timing + KV-cache model
  • Omniverse-rendered fleet twin
twinex · scenario
twinex.simulate({
  swap: "h100 → gb200",
  serving: "disaggregated",
  traffic: "2x"
}); // predicts p99 · $/tok
Scenario A

H100 · static · p99 240ms

Scenario B

GB200 · disagg · p99 150ms

Cost

$/tok −38%

Util

62% → 84%

02 / Explore

What-if scenarios, safely.

Model hardware mix, NVIDIA Dynamo disaggregation, batch sizes, MIG partitioning, autoscaling, and routing policy — then compare predicted utilization, latency, and cost side by side before you commit.

  • Hardware mix + Dynamo + MIG scenarios
  • Predicted p99, throughput, utilization
  • Validate a Voltra policy before rollout
03 / Decide

Plan capacity and budget with evidence.

Export predicted cost and capacity curves for finance and platform, with a target error band validated against a real run — so GPU purchases and topology changes are decisions, not bets.

  • Cost + capacity forecast curves
  • Prediction validated vs. real runs
  • Shareable plans for finance + platform
±6%
Forecast error
84%
Predicted util
-38%
Projected $/tok
By the numbers

Twinex, in production.

0
Forecast error band
0
Twin, whole fleet
0
Scenarios / minute
0
Blind prod changes

Simulate first. Decide second. Regret never.

Model your first hardware swap or topology change before it touches production. Start free, or let the team build a calibrated twin from your DCGM traces.