AI systems can generate, test, and select CUDA-kernel candidates, shifting some engineering work toward setting goals, checking results, and intervening when an agent stalls. The reported benchmark gains are tied to particular tasks, GPUs, and comparison methods—not a promise of faster performance on every workload.

AI can generate and test GPU kernels

CUDA is NVIDIA’s software platform for programming GPUs. A kernel is a small program that performs a task on a GPU. An AI-generated CUDA kernel is code produced to implement or optimize an operation in a framework such as PyTorch.

The process can involve generating several candidate kernels, testing their outputs and speed, then selecting a candidate that meets the task’s requirements. Stanford Scaling Intelligence Lab describes a search workflow that branches from optimization ideas into candidate implementations, checks correctness, and uses performance results to guide later rounds.

Engineers still set goals and check the results

In these workflows, engineers set the objective, review candidate outputs and performance, and step in when an agent gets stuck. Some generated code can be difficult to understand even when its correctness can be checked. Unusual bugs and more intensive review are also part of the reported engineering experience.

That makes verification a central part of the work. A benchmark can check whether output matches a reference under its own test conditions; it cannot establish that the code is safe or reliable in every deployment.

What the KernelBench results show

The Stanford Scaling Intelligence Lab reported selected KernelBench experiments on an NVIDIA L40S. Its results included performance equal to 101.3% of the PyTorch torch.matmul reference for a 4096×4096 matrix multiplication, and 484.4% of the torch.nn.LayerNorm reference for an input shape of (16, 64, 256, 256). A separate softmax result reached 111.8% of its reference for an input shape of (4096, 65536).

The experiments used a search with OpenAI o3 and Gemini 2.5 Pro over five rounds and 10 KernelBench Level 1 problems. The reference code used default FP32 precision; the benchmark allowed lower-precision solutions within a 1e-02 tolerance and checked numerical outputs across many random inputs. Those conditions matter: the figures describe selected tasks and input sizes, not arbitrary GPU code.

Gimlet Labs reported a separate experiment on an NVIDIA H100 using KernelBench Levels 1–3. Across its evaluated tasks, the company reported an approximately 1.8× geometric-mean speedup against the stronger result, for each task, from eager PyTorch or torch.compile. It excluded several cases where technically correct shortcuts produced disproportionate speedups. The result therefore applies to its evaluated set and comparison method.

The two experiments used different GPUs, task sets, and methods. Their headline figures are not a direct head-to-head comparison.

U.S. postings point to continued demand for CUDA skills

Reported Lightcast-based figures showed more U.S. job postings requiring CUDA skills in the first eight months of 2026 than in all of 2025. A September 2026 snapshot put NVIDIA’s active U.S. postings requiring CUDA skills at more than 300.

Together, the reported hiring figures and the engineering workflows point to a role that includes both generating candidates and judging them: people still define the task, assess results, and step in when the automated search stalls.