Onboarding the first endpoints.Talk to us
September 30, 2026justkyriecai

One night on GPU MODE: #2 on A100 matmul, #5 on conv2d

We pointed our RSI system at GPU MODE's public kernel leaderboards and let it run. No human wrote the submissions.

GPU MODE is where the GPU programming community competes. It grew out of the PyTorch world and runs open lectures, and its kernel competitions have been run with AMD on MI300 and MI355X, with NVIDIA on Blackwell, and on B200s hosted by Nebius. The leaderboards are public and scored on the organizers' own GPUs, so nobody grades their own homework. The people on them write kernels for a living; one of GPU MODE's founders, Mark Saroufim, sits at #8 on the same matmul board.

We gave two of these problems to our RSI system and let it run on its own. It wrote the code, submitted it, read the results and tried again.

Results

Leaderboard (A100) Ours Median entry Mean entry Rank
matmul_v2 (fp16 matrix multiply) 599 µs 696 µs (1.16x slower) 832 µs (1.39x slower) World Top 2
conv2d_v2 (2D convolution) 4.59 ms 18.8 ms (4.1x slower) 67.9 ms (14.8x slower) World Top 5

Live leaderboards: matmul_v2 · conv2d_v2 (A100 tab).

Median and mean over each participant's best entry: 39 people on matmul, 14 on conv2d. A few very slow entries pull the means up.

The matmul run took one night: 27 versions and 36 diagnostic experiments in about 14 hours. Its score equals #1's to the microsecond; the two differ by 0.02 nanoseconds, a floating-point rounding, and the leaderboard lists ours second. The conv2d result came from a single evening.

Why kernels decide how fast your model runs

A model is a chain of kernels. Every token your users wait for is a matrix multiply, then an attention kernel, then another matrix multiply, thousands of times over. The model's speed is the sum of those kernels' times, so a slow kernel is paid for on every token.

Take one request to a language model where matrix multiplies take, say, 70% of the time per token. Keep everything else the same and swap only the matmul kernel:

Matmul kernel Time per token Tokens per second
Ours 1.00 100
Median entry (1.16x slower) 1.11 90
Mean entry (1.39x slower) 1.27 79

Same model, same GPU, same engine; one kernel costs you 10–21% of your speed. For a vision model built on convolutions, where our conv2d is 4x faster than the median entry, the median kernel makes the whole model 2.5x slower.

Illustrative: time per token = 30% other work + 70% matmul × the kernel's slowdown; for the vision example, 50% convolution.

For a real, measured case, see how we made Qwen3.8-27B 11x faster on one GPU. One step there is a single new kernel: 1.9x faster than cuBLAS on its shapes, 6–8% faster end to end, with nothing else changed.

Search, not heroics

The pieces already exist: kernels, libraries, engines, quantization formats, settings. Making a model fast means finding the one path through them that is fastest on your hardware, in a space of thousands of combinations where most paths are dead ends and the measurements are noisy. People search that space slowly. An agent can walk it all night, measure every step and keep what works. On GPU MODE, one night was enough to tie the top matmul score.