Google and Nvidia have issued a joint announcement saying that users will be able to tap into up to a million Nvidia GPUs via newly launched A5X instances, as part of a push to cut inference costs and boost token throughput.
Google is pitching A5X as a purpose-built kit for agentic AI workloads, the sort where a gaggle of models nibble at a task piece by piece until something useful falls out.
They sit inside Google’s AI Hypercomputer portfolio, the same stack behind Gemini and the firm’s consumer and enterprise AI wares.
Google used the launch to roll out more Hypercomputer upgrades, including new virtual machines on custom Arm-based CPUs, native PyTorch TPU support and eighth-generation tensor processors.
A5X is the first Google instance line designed to run on Nvidia’s Vera Rubin AI GPUs, with the backend leaning on Nvidia network accelerators to span single and multi-cluster infrastructure.
On the plumbing side, A5X will use Nvidia ConnectX-9 NICs, built to accelerate AI workloads over Ethernet in cloud environments.
Paired with Google’s Virgo platform, that setup is designed to scale to 80,000 Rubin GPUs in a single cluster and 960,000 GPUs across a multisite cluster.
Virgo is Google’s way of tying together AI chips inside a single data centre, and it supports Google’s TPUs as well as Rubin.
Google claims Virgo can link 134,000 TPUs in one facility and more than 1 million chips across multiple sites, which is a lot of silicon to keep fed and cooled.
Nvidia says A5X delivers 10x lower inference costs per token and 10x higher throughput per megawatt than the previous generation.
Nvidia said Cadence and Siemens products run through its infrastructure and are available on Google Cloud.
Google’s Gemini platform is being lined up to deploy agentic models and workflows into sectors like cybersecurity, because nothing says calm and controlled like autonomous systems on a live network.







