Lemonade Fixes AMD APU Model Streaming, Drops OpenMOSS ROCm As ~40x Slower Than Vulkan
Lemonade's 2026.40 release candidate fixes model streaming on AMD APUs by sizing against the addressable GTT memory pool instead of the smaller fixed vRAM carve-out, resolving failures with models like DeepSeek-V4-Flash-IQ2XXS-DS4 on Ryzen AI Max (Strix Halo) hardware. The release also removes the OpenMOSS ROCm backend on Windows and Linux after benchmarks showed it running roughly 40x slower than the Vulkan backend, likely due to CPU fallbacks in the ROCm code path. A new launch agent for JetBrains' Junie is also added, alongside the prior 2026.39.1 stable release which introduced configurable vRAM auto-eviction and a new year-week versioning scheme.

