Onboarding the first endpoints.Talk to us
September 30, 2026justkyriecai

What inference optimization brings: Qwen3.8-27B from 27 to 302 tokens/s on one GPU

The same open model on the same RTX PRO 6000, one request at a time. Each step below is one optimization; together they make it 7–11x faster, with the same answers.

Most teams serve an open model the way its model card shows: official weights, default engine settings. That works, but it leaves most of the GPU's speed unused. We took Qwen3.8-27B on one RTX PRO 6000 (96 GB, $2 an hour to rent) and added one optimization at a time, measuring each step, to show what each layer of the inference stack is worth.

The workload is one request at a time: a coding agent, a chat window, anything where someone waits for every token. We measured median decode speed on three prompt sets: math (GSM8K), code (HumanEval) and chat (ShareGPT).

The ladder

Step Math Code Chat
0. Default: official BF16 weights 27 27 27
1. + Quantization (NVIDIA NVFP4) 67 67 67
2. + Speculative decoding (DFlash2) 270 213 164
3. + Our draft kernel 286 226 177
4. + Deeper quantization 302 252 196
Total speedup 11x 9x 7x

Tokens per second, median, one request at a time. Step 2 is SGLang's own best recipe for this card. Steps 0 and 1 have no speculation and run at the same speed on every prompt; they were measured on SGLang 0.5.17, the rest on SGLang 0.5.20.

Why each step helps

At batch 1, speed is bytes. Every decode step reads all the model's weights to produce a few tokens. The GPU's arithmetic sits mostly idle; what limits speed is how fast it can read memory, about 1.8 TB/s on this card. So there are only two levers: read fewer bytes per step, or get more tokens out of each step.

Step 1 reads fewer bytes. BF16 weights are 54 GB; NVFP4 is 22 GB. Speed goes up about 2.5x, close to the ratio of bytes. The answers do not measurably change.

Step 2 gets more tokens per step. A 2B draft model proposes 8 tokens, and the 27B model checks all of them in one pass, keeping about 3.6 on average. The draft's guesses are only accepted if the big model agrees, so the output distribution is unchanged. This is the biggest single jump, 2.5x to 4x, and it is more on math than on chat, because math is more predictable.

Steps 3 and 4 are where it gets specific. Once the recipe is in place, what is left is the time of one speculative step, 19.8 ms. We profiled it and went after the bytes it still reads. The draft model's large matrices take 3.2 GB per step in BF16, and cuBLAS already reads them near the memory's limit, so a better library call cannot help. We wrote a Triton kernel that reads them as int8: the draft runs 1.9x faster and the whole step drops to 18.7 ms. Then we quantized the layers NVIDIA's checkpoint keeps in 8-bit, taking the model from 21.9 GB to 17.1 GB and the step to 16.9 ms.

These last two steps are smaller, +6–8% and then another +6–12%. That is how inference optimization usually goes: public tools get you most of the way, and the last 10–20% takes a profiler, a kernel and a model-specific decision. For a service that runs all day, that last part is still 10–20% more tokens from the same GPUs.

Same answers

Greedy decoding, GSM8K (400 questions) and HumanEval (164 problems, tests executed):

GSM8K HumanEval
SGLang's recipe (step 2) 97.25% 90.85%
+ kernel (step 3) 97.25% 92.68%
+ deeper quantization (step 4) 97.25% 91.46%

Item by item, the differences are not significant. One cost does not show here: after step 4 the model writes about 21% more tokens on code, so a code task finishes no sooner even though each token is faster. On math and chat, tasks finish 15–16% sooner. If output length matters to you, stop at step 3.

Try it

Measured by us on one card, one run per configuration. If you serve open models and want to know what these steps are worth for yours, reach us at triody.ink.