Skip to content

#Open-source models

0 items today
10/6Tue
  1. DEV Community · MCP78

    FP8 pitfall: GPU bill dropped 47%, but the model outputs “!!!!!!”

    The author ran Qwen2.5 7B/32B/72B on a single AMD MI300X with vLLM ROCm, priced at $2.99/GPU-hr. The BF16 baseline was 7B at $0.227/M, 32B at $0.77/M, and 72B at $1.67/M output tokens.

    Why it matters: The author benchmarked FP8 quantization on the MI300X and found that per-token billing can hide the model's output degrading into gibberish, then gave a reusable way to verify it.

9/23Wed
5/7Thu
  1. Permission Protocol · AI Agent Incident Tracker80

    仿冒 OpenAI 仓库在 Hugging Face 登顶热门榜并获 24.4 万次下载后投递窃密木马

    Hugging Face 上一个仿冒 OpenAI Privacy Filter 的仓库登上热门榜第一、获得 244,000 次下载,随后在安装该模型的 Windows 机器上执行窃取凭据的 infostealer。

    Awaiting translation

    Why it matters: 复盘 Hugging Face 上仿冒 OpenAI 仓库的投毒链条,展示热门榜如何被当作信任信号利用。