by

Nvidia’s BlueField gives AI servers a kick up the backside

Networking outfit F5 has shown that Nvidia’s BlueField-3 DPUs can improve AI inference performance by up to 3.24x without adding more GPUs.

The results come from tests conducted at F5’s California laboratory, where ServeTheHome examined its BIG-IP Next for Kubernetes platform running on Nvidia hardware. The system moves traffic management and security operations onto dedicated data processing units, freeing host CPUs for other work.

F5’s Endpoint Picker technology examines GPU utilisation, available resources and cached data before directing incoming AI requests to the most appropriate processors. This improves on conventional round-robin load balancing, which distributes requests without considering what individual GPUs are doing.

The system can use previously processed data stored in the key-value cache, reducing the need for GPUs to repeat expensive calculations. It can route requests between large and smaller language models based on predefined policies, including token allowances and workload requirements.

F5 has integrated the technology with Nvidia’s inference infrastructure, and adapted its BIG-IP software to run on the Arm processors inside BlueField-3 DPUs. The arrangement handles network security, DDoS protection and application delivery without consuming valuable host CPU resources.

ServeTheHome examined a test cluster built around Supermicro servers equipped with eight Nvidia H100 accelerators. Additional systems handled Kubernetes infrastructure and storage, while BlueField-3 DPUs provided the networking and security functions.

Testing used the Qwen3-32B language model running at FP8 precision, with Nvidia’s AI Perf Tool generating workloads. F5 compared its DPU-based configuration against an Envoy AI Gateway running on host processors.

The tests covered different traffic patterns, including multi-turn conversations, mixed workloads and requests sharing cached information. Concurrency reached 200 simultaneous requests, with input sizes extending to 20,000 tokens.

Under lighter loads, the two configurations delivered similar performance. However, when demand exceeded the cluster’s available key-value cache capacity, F5’s system achieved up to 3.24 times the performance of the host-based configuration.

That figure represents the most demanding test scenario, and the comparison involved different hardware configurations. Nevertheless, the results suggest that smarter traffic routing could help operators squeeze more performance from expensive GPU clusters without throwing more accelerators at the problem.

The tests were conducted in F5’s laboratory, and ServeTheHome disclosed that the company sponsored its visit and accompanying video. Independent testing would establish whether the improvements hold across different models and GPU architectures.

F5 has incorporated BIG-IP Next for Kubernetes into Nvidia’s Common Networking Reference Architecture, targeting hyperscalers and enterprises running large AI clusters. F5 did not disclose pricing for the software or additional DPU hardware, leaving operators to decide whether the potential savings justify the investment.

 

TOPICS:
ai data centres  ·  AI inference  ·  BIG-IP Next for Kubernetes  ·  BlueField-3  ·  dpu  ·  F5  ·  GPU performance  ·  Nvidia  ·  nvidia h100

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews