FP8 量化致 Qwen2.5 输出异常与账单陷阱
Understand the story
2026-10-06,DEV Community 一篇 MCP 相关文章记录了作者在单张 AMD MI300X 上用 vLLM ROCm 版运行 Qwen2.5 7B/32B/72B 的实测:按 $2.99/GPU-hr 计价,BF16 基线下输出成本分别为 7B $0.227/M、32B $0.77/M、72B $1.67/M token。作者称改用 FP8 后 GPU 账单下降 47%,但模型开始输出“!!!!!!”这类异常内容,形成成本与质量的取舍陷阱。目前报道仅给出该单机实测的成本数字与异常现象,未提供更多复现细节或官方回应。
Generated by AI from reports · Updated 9 hours ago
Report timeline
Follow the reports to explore different perspectives.
- DEV Community · MCPSelectedFP8 pitfall: GPU bill dropped 47%, but the model outputs “!!!!!!”
The author ran Qwen2.5 7B/32B/72B on a single AMD MI300X with vLLM ROCm, priced at $2.99/GPU-hr. The BF16 baseline was 7B at $0.227/M, 32B at $0.77/M, and 72B at $1.67/M output tokens.
Interest in this story
Current interest 7·Peak within comparable coverage 9(Oct 6, 15:00)·Change over 24 hours within comparable coverage –
The trend compares the same participants under continuous, complete observation, so its coverage may be narrower than the current score. Hover or tap to view hourly interest; use the left and right arrow keys to navigate.