Intel Xeon 6 Granite Rapids Memory Scaling Performance From 6 To 12 MRDIMMs

Written by Michael Larabel in Memory on 17 February 2026 at 10:00 AM EST. Page 9 of 10. 1 Comment.
TensorFlow benchmark with settings of Device: CPU, Batch Size: 256, Model: ResNet-50. 12 x MRDIMM was the fastest.
TensorFlow benchmark with settings of Device: CPU, Batch Size: 512, Model: ResNet-50. 12 x MRDIMM was the fastest.

With TensorFlow wasn't much of an advantage at eight or ten MRDIMMs but popped back up with all twelve MRDIMMs.

OpenVINO benchmark with settings of Model: Person Detection FP16, Device: CPU. 12 x MRDIMM was the fastest.
OpenVINO benchmark with settings of Model: Person Detection FP16, Device: CPU. 12 x MRDIMM was the fastest.
OpenVINO benchmark with settings of Model: Face Detection FP16-INT8, Device: CPU. 12 x MRDIMM was the fastest.
OpenVINO benchmark with settings of Model: Face Detection FP16-INT8, Device: CPU. 12 x MRDIMM was the fastest.
OpenVINO benchmark with settings of Model: Face Detection Retail FP16-INT8, Device: CPU. 10 x MRDIMM was the fastest.
OpenVINO benchmark with settings of Model: Weld Porosity Detection FP16-INT8, Device: CPU. 12 x MRDIMM was the fastest.
OpenVINO benchmark with settings of Model: Noise Suppression Poconet-Like FP16, Device: CPU. 12 x MRDIMM was the fastest.
OpenVINO benchmark with settings of Model: Noise Suppression Poconet-Like FP16, Device: CPU. 12 x MRDIMM was the fastest.
OpenVINO benchmark with settings of Model: Age Gender Recognition Retail 0013 FP16-INT8, Device: CPU. 12 x MRDIMM was the fastest.
OpenVINO benchmark with settings of Model: Age Gender Recognition Retail 0013 FP16-INT8, Device: CPU. 12 x MRDIMM was the fastest.

The OpenVINO AI toolkit typically yielded worthwhile performance gains up through the twelve MRDIMM configuration.

Llama.cpp benchmark with settings of Backend: CPU BLAS, Model: Qwen3-8B-Q8_0, Test: Text Generation 128. 12 x MRDIMM was the fastest.
Llama.cpp benchmark with settings of Backend: CPU BLAS, Model: Qwen3-8B-Q8_0, Test: Prompt Processing 512. 6 x MRDIMM was the fastest.
Llama.cpp benchmark with settings of Backend: CPU BLAS, Model: Qwen3-8B-Q8_0, Test: Prompt Processing 2048. 6 x MRDIMM was the fastest.
Llama.cpp benchmark with settings of Backend: CPU BLAS, Model: gpt-oss-20b-Q8_0, Test: Prompt Processing 2048. 6 x MRDIMM was the fastest.
Llama.cpp benchmark with settings of Backend: CPU BLAS, Model: Llama-3.1-Tulu-3-8B-Q8_0, Test: Text Generation 128. 12 x MRDIMM was the fastest.
Llama.cpp benchmark with settings of Backend: CPU BLAS, Model: Llama-3.1-Tulu-3-8B-Q8_0, Test: Prompt Processing 2048. 6 x MRDIMM was the fastest.
Llama.cpp benchmark with settings of Backend: CPU BLAS, Model: Mistral-7B-Instruct-v0.3-Q8_0, Test: Text Generation 128. 12 x MRDIMM was the fastest.
Llama.cpp benchmark with settings of Backend: CPU BLAS, Model: DeepSeek-R1-Distill-Llama-8B-Q8_0, Test: Text Generation 128. 12 x MRDIMM was the fastest.
Llama.cpp benchmark with settings of Backend: CPU BLAS, Model: DeepSeek-R1-Distill-Llama-8B-Q8_0, Test: Text Generation 128. 12 x MRDIMM was the fastest.

Llama.cpp for text generation enjoyed all twelve MRDIMMs while for prompt processing the six MRDIMMs typically performed the best with at least the large language models tested.

OpenVINO GenAI benchmark with settings of Model: granite-3.0-8b-instruct, Device: CPU. 12 x MRDIMM was the fastest.
OpenVINO GenAI benchmark with settings of Model: TinyLlama-1.1B-Chat-v1.0, Device: CPU. 8 x MRDIMM was the fastest.
Related Articles