by

AMD introduces Instinct MI350P PCIe GPUs for enterprise AI deployments

AMD has unveiled its new Instinct MI350P PCIe accelerators, targeting enterprises looking to deploy AI workloads on-premises without requiring major upgrades to power, cooling, or rack infrastructure.

The MI350P comes in a dual-slot PCIe form factor designed for standard air-cooled servers, allowing organizations to add AI acceleration into existing systems rather than investing in dedicated GPU platforms. AMD positions the cards for inference workloads, retrieval-augmented generation (RAG), and AI models ranging from small deployments to larger enterprise-scale implementations.

On the hardware side, the Instinct MI350P is based on AMD’s CDNA 4 architecture with 128 Compute Units (CUs), or 8,192 Stream Processors, and 512 Matrix cores. It supports lower-precision formats such as MXFP4 and MXFP6 for higher throughput, alongside sparsity acceleration for INT8 and BF16 workloads. AMD claims up to 2,299 TFLOPS of performance and peak throughput reaching 4,600 TFLOPS at MXFP4 precision, paired with 144 GB of HBM3E memory delivering up to 4 TB/s bandwidth.

AMD was also quite keen to note software flexibility, with support for open AI frameworks including PyTorch, Kubernetes GPU Operator integration, and AMD’s own inference microservices. The company says its open-source enterprise AI stack is designed to simplify migration of existing workloads with minimal code changes while reducing licensing and operational costs.

The MI350P PCIe cards are aimed at organizations seeking to scale AI deployments within current data center environments, positioning them as a lower-barrier alternative to large accelerator platforms while still supporting modern AI inference and enterprise workloads.

 

TOPICS:
AMD  ·  instinct  ·  MI350P

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews