FP8 pitfall: GPU bill dropped 47%, but the model outputs “!!!!!!”
Original title: The FP8 trap: my GPU bill dropped 47% because the model was printing "!!!!!!"
The author ran Qwen2.5 7B/32B/72B on a single AMD MI300X with vLLM ROCm, priced at $2.99/GPU-hr. The BF16 baseline was 7B at $0.227/M, 32B at $0.77/M, and 72B at $1.67/M output tokens.
The author benchmarked FP8 quantization on the MI300X and found that per-token billing can hide the model's output degrading into gibberish, then gave a reusable way to verify it.
I had one hour on an AMD MI300X and one question: what does a token actually cost on it?
One GPU, vLLM's ROCm build, Qwen2.5 at 7B, 32B and 72B. Thirty-two requests in flight, 256 output tokens max, priced at the $2.99/GPU-hr list rate. Every number below is wall-clock time × that rate ÷ tokens produced.
The baselines came out clean. Four repeat checks of each untouched config, so I'd know how much the machine wobbles on its own:
- 7B: $0.227 per million output tokens
- 32B: $0.77
- 72B: $1.67, on a single GPU, which still surprises me a bit
The 72B repeats agreed to within 0.4%. That's a very quiet machine.
The obvious next move
FP8 halves the bytes per weight. vLLM will do it for you on load with one flag:
vllm serve Qwen/Qwen2.5-72B-Instruct --quantization fp8
Numbers:
- 32B: $0.77 → $0.449 (−41%)
- 72B: $1.67 → $0.872 (−47%)
Three checks each, all well outside the noise. If I'd stopped there I would have written a very different post, and I'd have been telling you to cut your bill in half.
The thing that didn't fit
I wasn't looking for a problem. I was reading the per-block lines and something was too neat.
At temperature 0, with the same 32 prompts, BF16 32B wrote about 6,900 output tokens per block. FP8 wrote 8,192 on every block. 8,192 is 32 × 256: every request ran all the way to the cap.
Models don't suddenly get more talkative because you changed the weight format. So I asked both servers the same six questions.
"Summarize the causes of World War I in about 80 words."
BF16: an 89-token answer about nationalism, alliances and an archduke.
FP8:
!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
…for all 256 tokens.
It was cheaper per token because it was producing junk tokens very quickly. The cost number was accurate. It was just measuring the wrong thing.
What was real
Same idea, done properly: checkpoints that were quantized ahead of time and published (RedHatAI's Qwen2.5-*-Instruct-FP8-dynamic). This time I checked the answers first.
They were right. 17 × 23 = 391. "Où est la gare ?" Sensible summaries. Output length matched BF16 within about 1% (72B: 6,361 vs 6,328 tokens per block).
Then the cost, three checks each:
| Model | BF16 | FP8 (pre-quantized) | change |
|---|---|---|---|
| 7B | $0.227/M | $0.185/M | −18% |
| 72B | $1.67/M | $1.13/M | −32% |
Noise bound under 3% on 7B and under 1% on 72B. So the real answer is about a third off a 72B on one MI300X. That's a smaller number than 47%, and you can actually use it.
The rule I'm keeping
A cost-per-token number on its own can't tell a cheaper model from a broken one. Before you trust a saving, check that the model is still doing the same work:
-
Output length. If average tokens per answer jumps, or suddenly everything hits
max_tokens, stop. Nothing else needed. - A handful of fixed prompts at temperature 0. Read them. Takes a minute.
- Repeat the baseline. If you don't know how much your server wobbles untouched, you can't tell a 5% win from Tuesday.
Number 1 is the cheap one, and the data for it is already sitting in your benchmark output. I'm building it into the tool I used here: if output length shifts between a baseline and a change, it says OUTPUT CHANGED instead of CHEAPER. That's coming in the next release.
Caveats, honestly
- One session, one workload: short prompts, 256 output tokens, concurrency 32. Long-context or long-output traffic will behave differently.
- On 72B, the FP8 server ran on the VM's second GPU and the baseline on the first. Same card model, not literally the same card.
- I only sampled the broken on-the-fly 72B by output length, not by reading answers. On-the-fly FP8 on 7B was fine (6 of 6 answers correct, −18%).
- Six prompts is a sanity check, not an accuracy benchmark. If FP8 matters to your product, run your own evals.
- Prices are the provider's list rate. Tokens and time are measured.
I don't know yet whether the on-the-fly FP8 behaviour is specific to this ROCm image or to larger Qwen2.5 models. If you've seen it, tell me.
Measured with Throttle, an open-source CLI that prices your inference in $/M tokens with confidence intervals and refuses to call a change a win inside your own noise. pipx install throttle-pro.
Source: DEV Community · MCP · dev.to