Skip to content

#Open-source models

1 item today
Today10/6Tue
  1. DEV Community · MCP78

    FP8 pitfall: GPU bill dropped 47%, but the model outputs “!!!!!!”

    The author ran Qwen2.5 7B/32B/72B on a single AMD MI300X with vLLM ROCm, priced at $2.99/GPU-hr. The BF16 baseline was 7B at $0.227/M, 32B at $0.77/M, and 72B at $1.67/M output tokens.

    Why it matters: The author benchmarked FP8 quantization on the MI300X and found that per-token billing can hide the model's output degrading into gibberish, then gave a reusable way to verify it.

10/5Mon
10/4Sun
  1. DEV Community · Cursor76

    Cursor ships Composer 2, and the API response strings give away its undisclosed Kimi K2.5 base

    On March 20, 2026, developer Fynn was debugging Cursor's OpenAI-compatible endpoint when the returned model ID came back as accounts/anysphere/models/kimi-k2p5-rl-0317-s515-fast — evidence that Composer 2 was post-trained with reinforcement learning on top of Moonshot AI's Kimi K2.5. The tweet hit 44.4 views within a day.

    Why it matters: One API debugging session ties together Cursor's undisclosed Kimi base, the licensing attribution dispute, and the cost landscape for Chinese versus U.S. models — a look at how the industry handles disclosure.

10/3Sat
10/2Fri
9/29Tue
9/25Fri
9/23Wed
9/22Tue
  1. Sebastian Raschka72

    小米发布开源权重模型 MiMo-V2.6-Pro,在 Artificial Analysis 智能指数上以 46 分成为开源权重模型第一,每任务成本 0.13 美元,输入 0.435 美元/1M tokens、输出 0.87 美元/1M tokens,采用 1.02T 总参数、42B 激活参数的 MoE 架构。

    Awaiting translation

    QuotedArtificial Analysis@ArtificialAnlys

    MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46). At $0.13 per Intelligence Index task, it lands on the Intelligence vs. Cost per Task Pareto frontier @Xiaomi has just released MiMo-V2.6-Pro, an open weights model with major advances in intelligence over its predecessor, MiMo-V2.5-Pro (Intelligence Index: 26). Despite the improvement, it retains the same attractive pricing at $0.435 per 1M input tokens (with a 99% cache-hit discount) and $0.87 per 1M output tokens. This makes MiMo-V2.6-Pro one of the most cost-efficient models to deploy. MiMo-V2.6-Pro is an MoE model with 1.02T total parameters and 42B active parameters. Stay tuned for additional analysis of the model. Check out MiMo-V2.6-Pro full benchmarking breakdown here: https://artificialanalysis.ai

9/16Wed
9/15Tue
  1. Cline · Blog62

    Cline releases the open-source desktop app Cline Desktop, aimed at open-weight models.

    Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.

    Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.

9/12Sat
9/9Wed
9/3Thu
  1. Cline · Blog74

    Cline 如何把 1100 万用户迁移到最大一次 harness 升级

    Cline 把 VS Code 扩展从约 76,000 行单体核心迁移到 Cline SDK,并自建灰度发布机制:一个安装包内打包 loader、legacy 和 next 两套扩展,由 PostHog 功能开关按百分比决定激活哪套,崩溃时自动回退到 legacy,开关可随时降到 0% 作为 kill switch。

    Awaiting translation

    Why it matters: Cline 官方复盘如何把 1100 万用户的 VS Code 扩展迁到新 harness,含灰度机制与前后指标对比。

8/29Sat
  1. Martin Alderson38

    GLM-5.3 Flash 跑在国产硬件上意味着什么

    Z.AI 确认 GLM-5.3 Flash 的全部推理运行在国产硬件上,但未公布芯片厂商、吞吐或功耗数据,也无人独立验证。作者推测其使用的是华为昇腾 910c 系列,该芯片约 600W、INT8 算力约 1.6PFLOP/s,性能约为四年前 H100 的 60%,且缺少原生 FP8 支持。作者认为真正的瓶颈是缺乏可量产的 EUV 光刻,业界普遍认为 2030 年前难以实现规模化生产。

    Awaiting translation

8/27Thu
  1. Cline · Blog71

    Cline 实测八个模型做 IMO 2026:DeepSeek V4 Flash 以 0.12 美元拿到金牌线

    Cline 让八个模型在自家 harness 里做 IMO 2026 六道题,证明由 GPT-5.5 和 Claude Opus 5 双盲按 0–7 分制评分、Gemini 3.1 Pro 仲裁,金牌线为 29 分。

    Awaiting translation

    Why it matters: Cline 用同一套 harness 盲评八个模型做 IMO 2026,给出分数与单次成本对照,可看开源权重模型的实际性价比。

8/23Sun
  1. Martin Alderson62

    开源权重模型的夏天:推理定价战与算力约束如何改变前沿实验室的处境

    作者认为这个夏天是开源权重模型的转折点,多数智能体任务已不再必须依赖前沿模型。他列举了 OpenAI 将 5.6 Luna 降价 80%、Sol 降价 20%,Meta 在贡献者档位把 Muse Spark 1.2 压到 $0.10/$0.20 per MTok,以及 Anthropic 因算力紧张而难以跟进降价。

    Awaiting translation

8/22Sat
8/19Wed
  1. Cline · Blog71

    Cline's Evaluation Methodology and Trace for Open-Weight Models

    Cline has open-sourced its evaluation methodology for open-weight models, along with a hill-climbing score and trace worth over one thousand dollars, available for download and analysis. The post lays out five heuristics from the Hill Climber's Checklist: set a North Star metric, quantify noise, break down failure modes by task/model/vendor, don't assume more thinking is always better, and keep a private evaluation set.

    Why it matters: Cline shares its evaluation methodology and a trace worth over a thousand dollars, offering five transferable hill-climbing heuristics that teams building their own harness can reference.

8/17Mon
8/11Tue
  1. Sean Goedecke · Blog52

    为什么本地模型不会胜出:批处理与 GPU 效率决定推理成本

    Sean Goedecke 认为本地模型不会成为主流,绝大多数推理仍会发生在 AI 数据中心。他给出的理由是:前沿模型体积远超消费级设备,而数据中心可通过批处理数百名用户的请求摊薄成本,加上 B200 相比 RTX 4090 在同等功耗下约有 3 倍 flops 和近 4 倍内存带宽,本地运行约需多消耗 30 倍资源;家庭推理的电力成本约每月 50 至 300 美元。

    Awaiting translation

8/5Wed
  1. Chen Dahuang · AI Coding 实录66

    DeepSeek V4 Flash 正式版深度体验:便宜、快、1M 上下文、内置搜索

    作者深度体验几天 DeepSeek V4 Flash 0731 正式版后总结:便宜到跑批处理、Agent 循环和几十轮对话账单基本无感,速度快到配合 Agent 工具循环每步几秒内完成,1M 上下文可容纳整个仓库和完整对话历史、无需频繁 compact,结合 Cache 打折长上下文成本还能再降。

    Awaiting translation

8/2Sun
7/7Tue
7/6Mon
6/28Sun
6/22Mon
6/15Mon
6/7Sun
  1. Thorsten Ball · Register Spill12

    Joy & Curiosity #89:Anthropic《When AI builds itself》、Ted Chiang 论 AI 意识与 Ladybird 停收公开 PR

    Thorsten Ball 的 Joy & Curiosity 通讯在写了三年后进入两三周、最多四周的暑期休更。本期链接包括 Anthropic 发布的《When AI builds itself》文档、Ted Chiang 关于 AI 是否具有意识的文章,以及 Ladybird 浏览器宣布不再接受公开 pull request,理由是 AI 工具已迅速改变开源信任的经济学。

    Awaiting translation

5/23Sat
5/7Thu
  1. Permission Protocol · AI Agent Incident Tracker80

    仿冒 OpenAI 仓库在 Hugging Face 登顶热门榜并获 24.4 万次下载后投递窃密木马

    Hugging Face 上一个仿冒 OpenAI Privacy Filter 的仓库登上热门榜第一、获得 244,000 次下载,随后在安装该模型的 Windows 机器上执行窃取凭据的 infostealer。

    Awaiting translation

    Why it matters: 复盘 Hugging Face 上仿冒 OpenAI 仓库的投毒链条,展示热门榜如何被当作信任信号利用。

5/6Wed
5/4Mon
4/30Thu