NVIDIA software stack cuts token costs on Blackwell GPUs

NVIDIA · June 30, 2026 · ✓ verified

NVIDIA announces its inference software stack reduces token cost and increases throughput on Blackwell GPUs.

  • Main announcement: NVIDIA’s full-stack inference software for the Blackwell platform has reduced token costs by up to 5x on the DeepSeek V4 model in about one month, and combining system-level optimizations (disaggregated serving, large expert parallelism, NVFP4, multi-token prediction) can increase throughput by up to 20x. The blog cites partner results such as Baseten reporting up to 50% more tokens/sec, DigitalOcean / Hippocratic AI reporting ~30% higher inference throughput while maintaining sub-half-second time-to-first-response, and day-zero deployment recipes for vLLM and SGLang.
  • Background and implementation details: The announcement explains the stack connects three layers — Production Operation, Application Acceleration, and Infrastructure Access — and leverages open source frameworks (PyTorch, vLLM, SGLang) and runtimes (TensorRT-LLM, NVIDIA Dynamo). It references concrete software features and optimizations (DFlash speculative decode up to 15x throughput, NVLink interconnect, NVFP4 precision) and notes these improvements were observed in production-focused tests and partner deployments within a short timeframe (about a month).
Keep reading
Video explains everyday dependence on data centers DataBank · Jul 21 Arista launches AI-driven zero trust branch security Arista Networks · Jul 20 NVIDIA showcases AI and graphics breakthroughs at SIGGRAPH NVIDIA · Jul 20 Japan deepens U.S. ties with $550 billion investment Jefferies.com · Jul 20
Telborg · US Data Centers
Track the US data-center buildout — every day.

Real-time verified news and daily AI-written briefings, built from primary sources — power, grid, permits, land, financing. Start free.

Get Telborg Pro · $189/mo Get the daily briefing — free →

Every field traced to a primary source.