AMD PerfOpt is a planned Linux 7.4 optimization for qualifying Radeon integrated GPUs, with reported local AI inference gains that differed across tested systems. Results published September 29, 2026, put typical gains at 2–4% on one desktop and as high as 23% on a laptop.
What AMD PerfOpt changes in Linux 7.4
The PerfOpt patch was reported queued for the Linux 7.4 IOMMU development tree on September 28, 2026. The planned AMDGPU behavior is to enable it by default for qualifying integrated GPUs that are already in identity mode. A user can disable the feature with the kernel parameter amdgpu.iommu_perfopt=0.
How PerfOpt changes IOMMU access
An input-output memory management unit (IOMMU) translates device memory addresses and enforces access permissions. In the described identity-mode case, PerfOpt lets a qualifying, privileged integrated GPU access system memory without that address-translation step. The IOMMU still enforces its read/write permission bits.
The optimization is intended for integrated devices such as GPUs, not discrete graphics cards or other external hardware. It concerns the GPU, rather than an NPU.
Reported results on three tested systems
The tests compared PerfOpt enabled and disabled on the same development kernel, using Lemonade with a llama.cpp backend for local AI inference. The reported results varied by system:
| Tested system | Reported result |
| Framework Desktop with AMD Ryzen AI Max+ 395 (Strix Halo), 64 GB LPDDR5-8000 | Typically 2–4% better performance |
| ASUS Zenbook S16 with AMD Ryzen AI 9 365 (Strix Point), 16 GB LPDDR5-7500 | Gains as high as 23% |
| Laptop with AMD Ryzen AI 5 340 (Krackan Point) | Performance gains and lower time to first token |
The first two figures describe particular systems and local inference tests: 2–4% is the typical result for the tested Strix Halo desktop, while 23% was a maximum in tests on the Strix Point laptop. The Krackan Point result was described without a percentage. Detailed test scenarios included Vulkan and AMD ROCm backends.
These measurements concern local inference on the three tested integrated-GPU systems. They do not establish the same performance change for every Radeon GPU or workload.