
Lemonade 11.9 introduces experimental AMD ROCm HRX backend
Lemonade 11.9 arrives as the newest feature release of the open-source local AI server, now supporting an experimental ROCm HRX backend for Llama.cpp. The HRX system provides a lightweight alternative to the traditional HIP implementation and targets Radeon RX 7900 series GPUs as well as Strix Halo APUs.
AMD engineer Stella Laurenzo explains that HRX is a focused subset of ROCm designed for client-side workloads, allowing faster code generation into AMDGPU assembly. Initial benchmarks show a 30 to 50% token-per-second uplift on prefill tasks and about a 10% gain on decode operations.
The release also adds Qwen3-Next support to the pinned Llama.cpp backend and includes other stability improvements. Lemonade’s core promise of “100% free and private” AI continues, now with the possibility of higher performance on AMD hardware without needing datacenter-grade stacks. Users can test the HRX backend on Linux, while future expansions to Windows and macOS are planned. Meer details staan op de pagina van Phoronix.









