Skip to content
Trending storyUpdating

FP8 量化致 Qwen2.5 输出异常与账单陷阱

1 report1 reporting sourceUpdated 13 hours ago

Understand the story

AI synthesis

2026-10-06,DEV Community 一篇 MCP 相关文章记录了作者在单张 AMD MI300X 上用 vLLM ROCm 版运行 Qwen2.5 7B/32B/72B 的实测:按 $2.99/GPU-hr 计价,BF16 基线下输出成本分别为 7B $0.227/M、32B $0.77/M、72B $1.67/M token。作者称改用 FP8 后 GPU 账单下降 47%,但模型开始输出“!!!!!!”这类异常内容,形成成本与质量的取舍陷阱。目前报道仅给出该单机实测的成本数字与异常现象,未提供更多复现细节或官方回应。

Generated by AI from reports · Updated 9 hours ago

Report timeline

Follow the reports to explore different perspectives.

10/6
  1. DEV Community · MCPSelected
    FP8 pitfall: GPU bill dropped 47%, but the model outputs “!!!!!!”

    The author ran Qwen2.5 7B/32B/72B on a single AMD MI300X with vLLM ROCm, priced at $2.99/GPU-hr. The BF16 baseline was 7B at $0.227/M, 32B at $0.77/M, and 72B at $1.67/M output tokens.

Interest in this story

Current interest 7·Peak within comparable coverage 9(Oct 6, 15:00)·Change over 24 hours within comparable coverage –

02.557.51010/615:0010/618:0010/621:0010/700:00

The trend compares the same participants under continuous, complete observation, so its coverage may be narrower than the current score. Hover or tap to view hourly interest; use the left and right arrow keys to navigate.