09/16/2026 | Press release | Distributed by Public on 09/16/2026 09:25
Why It Matters: As AI inference moves into production, customers need platforms that can support larger and more diverse models while realizing performance improvements, flexibility and value over time. Intel's MLPerf Inference v6.1 results prove how software optimization can extend the useful performance of deployed infrastructure, helping customers get more from systems they already use while scaling across CPU and GPU compute.
Customer and partner participation on Intel platforms also increased from 29 results in v6.0 to 39 in v6.1, expanding third-party validation beyond Intel's own submissions². Oracle contributed its first Intel-based submission; Red Hat delivered its first Xeon CPU inference submission, and Quanta Cloud Technology and Supermicro provided the first partner submissions using Intel Arc Pro B70.
About Intel Xeon 6 Results: Intel broadened its Xeon participation in MLPerf v6.1 from two benchmarked Xeon 6 SKUs in v6.0 to five, increasing the number of CPU inference results from 24 to 35. Intel Xeon remains the only standalone server CPU represented in MLPerf Inference submissions.
About Intel Arc Pro B70 Results: Intel Arc Pro B70 submissions expand on how software improvements can extend performance across a range of AI models. A single node configured with four Intel Arc Pro B70 GPUs provides 128GB of VRAM and supported submissions across Llama 3.1 8B, Llama 2 70B, gpt-oss-120B, Whisper and end-to-end retrieval-augmented generation (E2E-RAG). On the same four-GPU system used in v6.0, gpt-oss-120B Server performance improved 36%, while Offline performance increased 27%, reflecting continued maturity of Intel's software stack³.
About the New E2E-RAG Benchmark: Intel also co-developed and submitted results for the new MLPerf Inference v6.1 end-to-end retrieval-augmented generation benchmark, one of the round's most complex new workloads. In a single measured run on a system combining an Intel Xeon 6787P processor with four Intel Arc Pro B70 GPUs, the workload was split across the CPU and GPUs, with each handling different stages of the AI pipeline: Xeon handles embedding, reranking, vector search and Small Language Model (SLM), while Arc Pro B70 GPUs perform Large Language Model (LLM) generation. The result demonstrates how optimized CPU and GPU compute can work together across a complete AI workflow.
Intel's ongoing optimization work extends beyond benchmark results. Xeon improvements are upstreamed into widely used AI frameworks so customers can benefit from software advances on deployed infrastructure, while optimization work on Intel Arc Pro B70 helps advance the kernels, frameworks, and serving stack for future Intel GPU products.
More Context: MLPerf Inference v6.1 results