NVIDIA says Vera Rubin boosts agentic AI efficiency
NVIDIA
· August 24, 2026
· ✓ verified
NVIDIA has published a blog post explaining why agentic AI drives much higher token demand and claiming new efficiency gains for Vera Rubin NVL72 systems.
- The post says agentic workflows keep reasoning across steps, querying databases, invoking sub-agents, and synthesizing results, which causes context to accumulate and raises token usage.
- NVIDIA says its measured data shows up to 30x higher throughput per megawatt for Vera Rubin NVL72 versus GB300 NVL72 on agentic workloads, using the SemiAnalysis AgentX workload; it also says Vera Rubin can reach up to 35x lower cost per million tokens.
- The article is a commentary/marketing-style technical blog rather than a standalone commercial deal announcement, and it references prior NVIDIA product performance claims and related software/hardware optimizations.