May 24, 2026
NVIDIA Dynamo disaggregated serving GA
Disaggregated prefill/decode via NVIDIA Dynamo on H100/H200 clusters. 40% reduction in first-token latency. NIM-compatible API. Visual serving graph GA.
A new release every Friday. Subscribe to the RSS feed or follow the team on X.
May 24, 2026
Disaggregated prefill/decode via NVIDIA Dynamo on H100/H200 clusters. 40% reduction in first-token latency. NIM-compatible API. Visual serving graph GA.
May 17, 2026
Per-model content safety policies · GPU budget caps per inference route · escalation routing by org chart.
May 10, 2026
GPU utilization, memory bandwidth, NVLink throughput now stream into Voltra dashboards. Correlate GPU metrics with inference latency.
May 3, 2026
Block deployments that regress inference throughput or p99 latency. GitHub Action included.
Apr 26, 2026
Discover and install NIM-optimized models from inside the serving graph builder.
Apr 19, 2026
Native exporters for Splunk, Datadog, Snowflake — with per-model GPU cost breakdown.