AMD ROCm vs. Vulkan Performance For Lemonade Local AI Server With Llama.cpp

Written by Michael Larabel in Display Drivers on 5 October 2026 at 09:50 AM EDT. Page 1 of 3. 8 Comments.

A Phoronix Premium reader recently inquired about seeing some new benchmarks of AI performance between Vulkan and AMD ROCm back-ends to see how the situation plays out today. Here are some fresh benchmarks on two separate systems with AMD Radeon graphics and using the Lemonade local AI server while exploring the difference between the ROCm and Vulkan Llama.cpp back-ends.

Lemonade local AI server

It's been nearly a year since last looking at the RADV Vulkan vs. AMD ROCm performance with Llama.cpp on AMD RDNA4. Going back then there were a number of models where using the Llama.cpp Vulkan back-end was faster than using AMD ROCm of the time.

Lemonade on Ubuntu Linux with AMD Radeon hardware

With today's testing the Lemonade 2026.39.1 local AI server binary was used for testing for easy reproducibility with its built-in benchmarking support. Lemonade 2026.39.1 uses Llama.cpp b10825.

ROCm vs. Vulkan Lemonade Strix Halo

Two systems were used for this fresh ROCm vs. Vulkan testing, one a System76 Thelio Major workstation with the AMD Radeon AI PRO R9700 and the other the Framework Desktop with the AMD Ryzen AI Max+ 395 Strix Halo. Completely different hardware and slightly different software stacks to offer distinct looks at Vulkan vs. ROCm with Lemonade/Llama.cpp. Thanks to Phoronix Premium supporters that help make all of the daily Linux benchmarking possible at Phoronix 22+ years and going.

Radeon AI PRO R9700 Lemonade
Related Articles