What inference optimization brings: Qwen3.8-27B from 27 to 302 tokens/s on one GPU
The same open model on the same RTX PRO 6000, one request at a time. Each step below is one optimization; together they make it 7–11x faster, with the same answers.
Measurements, methods, and the numbers we had to take back.
What inference optimization brings: Qwen3.8-27B from 27 to 302 tokens/s on one GPU
The same open model on the same RTX PRO 6000, one request at a time. Each step below is one optimization; together they make it 7–11x faster, with the same answers.
One night on GPU MODE: #2 on A100 matmul, #5 on conv2d
We pointed our RSI system at GPU MODE's public kernel leaderboards and let it run. No human wrote the submissions.