AMD ROCm vs. Vulkan Performance For Lemonade Local AI Server With Llama.cpp

Written by Michael Larabel in Display Drivers on 5 October 2026 at 09:50 AM EDT. Page 3 of 3. 9 Comments.

And then firing up the System76 Thelio Major with Ryzen Threadripper and AMD Radeon AI PRO R9700 graphics, to see if it was the same story as the Strix Halo APU or different.

Lemonade benchmark with settings of Backend Comparison (Scenario: code-debug, Model: Qwen3-14B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-debug, Model: Qwen3-14B-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-short, Model: Qwen3-14B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-short, Model: Qwen3-14B-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-debug, Model: Qwen3.5-4B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-debug, Model: Qwen3.5-4B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-short, Model: Qwen3.5-4B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-short, Model: Qwen3.5-4B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-debug, Model: MiniCPM4-8B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-debug, Model: MiniCPM4-8B-GGUF). AMD ROCm was the fastest.

Even with the completely different hardware and different Ubuntu release and kernel, it was the same story as seen on the Framework Desktop: the Vulkan back-end was delivering more tokens per second while the ROCm back-end enjoyed lower latency.

Lemonade benchmark with settings of Backend Comparison (Scenario: code-explain, Model: Qwen3-14B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-explain, Model: Qwen3-14B-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-short, Model: MiniCPM4-8B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-short, Model: MiniCPM4-8B-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-explain, Model: Qwen3.5-4B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-explain, Model: Qwen3.5-4B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-explain, Model: MiniCPM4-8B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-explain, Model: MiniCPM4-8B-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: chat-long-output, Model: Qwen3-14B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: chat-long-output, Model: Qwen3-14B-GGUF). AMD ROCm was the fastest.

The time to first token with ROCm tended to be quite significant while the Vulkan throughput advantage was typically more incremental . In a few cases the Vulkan back-end with the Radeon AI PRO R9700 did enjoy faster times to first token compared to the ROCm back-end.

Lemonade benchmark with settings of Backend Comparison (Scenario: chat-long-output, Model: Qwen3.5-4B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: chat-long-output, Model: Qwen3.5-4B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-debug, Model: Qwen3-Coder-Next-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-debug, Model: Qwen3-Coder-Next-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-short, Model: Qwen3-Coder-Next-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-short, Model: Qwen3-Coder-Next-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: chat-long-output, Model: MiniCPM4-8B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: chat-long-output, Model: MiniCPM4-8B-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-debug, Model: DeepSeek-Qwen3-8B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-debug, Model: DeepSeek-Qwen3-8B-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-short, Model: DeepSeek-Qwen3-8B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-short, Model: DeepSeek-Qwen3-8B-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-explain, Model: Qwen3-Coder-Next-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-explain, Model: Qwen3-Coder-Next-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-explain, Model: DeepSeek-Qwen3-8B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: code-explain, Model: DeepSeek-Qwen3-8B-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: chat-long-output, Model: Qwen3-Coder-Next-GGUF). AMD ROCm was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: chat-long-output, Model: Qwen3-Coder-Next-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: chat-long-output, Model: DeepSeek-Qwen3-8B-GGUF). Vulkan was the fastest.
Lemonade benchmark with settings of Backend Comparison (Scenario: chat-long-output, Model: DeepSeek-Qwen3-8B-GGUF). AMD ROCm was the fastest.

So for those wondering, the Vulkan back-end of Llama.cpp is still working quite well overall and working across GPU vendors. The AMD ROCm back-end though as shown on both the Radeon AI PRO R9700 and Strix Halo Radeon 8060S Graphics was providing lower latency performance.

If you enjoyed this article consider joining Phoronix Premium to view this site ad-free, multi-page articles on a single page, and other benefits. PayPal or Stripe tips are also graciously accepted. Thanks for your support.

Related Articles
About The Author

Michael Larabel is the principal author of Phoronix.com and founded the site in 2004 with a focus on enriching the Linux hardware experience. Michael has written more than 20,000 articles covering the state of Linux hardware support, Linux performance, graphics drivers, and other topics. Michael is also the lead developer of the Phoronix Test Suite, Phoromatic, and OpenBenchmarking.org automated benchmarking software. He can be followed via Twitter, LinkedIn, or contacted via MichaelLarabel.com.