Results published October 9, 2026, put AMD PerfOpt’s commonly observed gains at 4–12% across AI benchmarks on a Lenovo ThinkPad P14s Gen 4 with an AMD Ryzen 7 PRO 7840U and AMD Radeon 780M. The tests used Ubuntu 26.04 LTS and a kernel build from the IOMMU next Git tree, with PerfOpt toggled as the sole reported system change.
The range applies to those varied benchmarks on that laptop configuration. This 7840U result adds a different data point to earlier coverage of PerfOpt results across other Ryzen systems.
Lemonade latency and llama.cpp workloads
In Lemonade, the most noticeable reported improvement was lower time to first token—the delay before a local AI model begins responding. Throughput also improved, but less. The tests covered code-debug, code-short, code-explain and chat-long-output scenarios, using Vulkan and AMD ROCm backends. Model captions included Qwen3-14B-GGUF, Qwen3.5-4B-GGUF, MiniCPM4-8B-GGUF, DeepSeek-Qwen3-8B-GGUF and Qwen3-Coder-Next-GGUF.
The tests also reported improvements in llama.cpp’s Vulkan token generation and prompt processing. Selected generation runs used 128-token prompts, while prompt-processing cases used 2,048 tokens. Models in those cases included gpt-oss-20b-Q8_0, Qwen3.5-9B-Q8_0, GLM-4.7-Flash-IQ4_XS and Llama-3.1-Tulu-3-8B-Q8_0. The overall 4–12% range describes the range reported across various AI benchmarks, not a result assigned to any one model or backend.
The test system and PerfOpt toggle
The benchmark setup paired the ThinkPad P14s Gen 4’s Ryzen 7 PRO 7840U with its Radeon 780M integrated graphics. It ran Ubuntu 26.04 LTS on a kernel build from the IOMMU next Git tree. PerfOpt was toggled between runs; no other system changes were reported.
The Linux 7.4 plan and proposed IOMMU trade-off
An IOMMU, or input-output memory management unit, manages how devices access system memory. AMD’s PerfOpt is an IOMMU-related optimization intended to reduce overhead for compatible integrated graphics.
On September 28, AMD engineer Mario Limonciello posted a patch proposing PerfOpt for eligible integrated GPUs in identity mode, where the GPU uses direct memory mappings. After reviewer Jason Gunthorpe raised design questions, Limonciello said he would revise the proposal. Its documentation describes a trade-off: lower DMA latency in exchange for disabling ATS, PRI, PASID and SVA for the GPU, and removing IOMMU DMA containment for it.
The October 9 report described default activation with Linux 7.4 as planned.