Changelog

Shipped this week.

A new release every Friday. Subscribe to the RSS feed or follow the team on X.

v2026.5.0

May 24, 2026

NVIDIA Dynamo disaggregated serving GA

Disaggregated prefill/decode via NVIDIA Dynamo on H100/H200 clusters. 40% reduction in first-token latency. NIM-compatible API. Visual serving graph GA.

VoltraGPUNIM
v2026.4.3

May 17, 2026

NeMo Guardrails v1.2 integration

Per-model content safety policies · GPU budget caps per inference route · escalation routing by org chart.

PolicyGuardrails
v2026.4.2

May 10, 2026

DCGM GPU telemetry in Voltra

GPU utilization, memory bandwidth, NVLink throughput now stream into Voltra dashboards. Correlate GPU metrics with inference latency.

ObservabilityGPU
v2026.4.1

May 3, 2026

TensorRT-LLM compilation pipeline

Block deployments that regress inference throughput or p99 latency. GitHub Action included.

EvalsDX
v2026.4.0

Apr 26, 2026

NVIDIA NIM microservice registry

Discover and install NIM-optimized models from inside the serving graph builder.

InferenceMarketplace
v2026.3.4

Apr 19, 2026

SIEM export with GPU cost attribution

Native exporters for Splunk, Datadog, Snowflake — with per-model GPU cost breakdown.

SecurityObservability