Nvidia doubles down on AI factories as it showcases massive Vera Rubin performance gains Artificial intelligence chip king Nvidia Corp. today revealed a fresh trove of performance benchmarks and architectural milestones for its next-generation Vera Rubin platform as it edges closer to global availability. The new numbers are impressive, but on a higher level they also underscore the potency of Nvidia’s approach that combines custom silicon with its broader hardware ecosystem and aggressive software optimization to squeeze even greater efficiency gains for AI workloads. The latest milestones showcase the effort the chipmaker has made toward fine-tuning its ecosystem to tackle the increased performance demands and exploding infrastructure costs of next-generation “agentic” AI workloads. Impressive gains One of the key revelations came from Nvidia’s partner CoreWeave Inc., the neocloud company that provides rented, cloud-based access to its graphics processing unit platforms. In early production runs, CoreWeave showed that it was able to squeeze out an incredible 10 times more tokens per watt when running DeepSeek Ltd.’s R1 model on the new Vera Rubin platform, compared with the previous generation GB200 NVL72 system based on Blackwell. It’s an impressive leap, but CoreWeave showed that the optimizations made to Vera Rubin can also be applied to the GB200 NVL72 system as well, improving throughput per megawatt by more than four times over a three-month span. Nvidia said these gains were validated across 250,000-plus distinct configurations and more than 1.4 million GPU hours of testing. Nvidia also wheeled out some new numbers for its Vera central processing units for running autonomous AI agents that can perform work on behalf of humans with minimal supervision. The Vera CPUs were built on a custom microarchitecture specification code-named Olympus core, and are highly optimized to process the irregular control flows of AI agents. In the latest benchmarks, Nvidia demonstrated 1.9 times faster agentic performance and a six-fold improvement in latency over x86-based alternatives. The company also showed that Vera surpassed Advanced Micro Devices Inc.’s flagship EPYC Turin CPU by almost 100% on selected industry benchmarks, underscoring the importance of custom silicon in eliminating CPU-side bottlenecks. Infra optimizations There’s no doubt that Nvidia’s processors are among the best in the business, but the company made it clear that these gains could only be achieved thanks to its strategy that’s focused on hardware and software co-design. The company also develops the required software and networking systems needed to optimize the performance of its chips so as to maximize power and cooling efficiency. Nvidia revealed that by dynamically optimizing the full infrastructure and energy stack, it could deploy 40% more GPUs within the same power envelope. At the same time, it showed how it can also minimize the environmental impact of its new processors by cooling them in a 45°C closed-loop liquid-cooled system that saves about 4 million gallons of water per megawatt annually over standard cooling methods. The company also talked about its latest generation of networking fabrics, showing how they continue to outperform the best generic Ethernet architectures. The sixth-generation NVLink 6 interconnect delivered 2.3 times higher simulated decode throughput for massive large language models than Ethernet-based networks. Meanwhile, the company’s Spectrum-X platform enabled 1.6 times faster remote direct memory access bandwidth with 1.7 times fewer switches. This integrated network results in five times greater optical power efficiency and 10 times more reliability, Nvidia said. In particular, the company revealed, its latest Spectrum-X platform, Spectrum-6, is arriving at AI factories from the likes of CoreWeave, Microsoft Corp., Nebius B.V., SpaceXAI Corp. and Tesla Inc. Nvidia is currently racing to ship the new Vera Rubin systems to customers and partners, including the likes of Google Cloud, Microsoft Azure, Meta Platforms Inc., Oracle Cloud Infrastructure, Dell Technologies Inc., OpenAI Group PBC and CoreWeave. As those platforms inch closer to general availability, the release of the latest benchmarks is clearly a strategic move by Nvidia, aiming to show that it now has all of the pieces in the puzzle for its customers to build the “AI factories” of the future. With reporting from Robert Hof Images: Nvidia A message from John Furrier, co-founder of SiliconANGLE: Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities. 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network. Are you AWS customer? Support SiliconANGLE Financially by buying your AWS services from our Marketplace portal page and links. About SiliconANGLE Media SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI. Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.
Nvidia doubles down on AI factories as it showcases massive Vera Rubin performance gains
Full Article
Original Source
Read the full article at Siliconangle →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.