Skip to content

#Open-source models

0 items today
10/2Fri
9/22Tue
  1. Sebastian Raschka72

    小米发布开源权重模型 MiMo-V2.6-Pro,在 Artificial Analysis 智能指数上以 46 分成为开源权重模型第一,每任务成本 0.13 美元,输入 0.435 美元/1M tokens、输出 0.87 美元/1M tokens,采用 1.02T 总参数、42B 激活参数的 MoE 架构。

    Awaiting translation

    QuotedArtificial Analysis@ArtificialAnlys

    MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46). At $0.13 per Intelligence Index task, it lands on the Intelligence vs. Cost per Task Pareto frontier @Xiaomi has just released MiMo-V2.6-Pro, an open weights model with major advances in intelligence over its predecessor, MiMo-V2.5-Pro (Intelligence Index: 26). Despite the improvement, it retains the same attractive pricing at $0.435 per 1M input tokens (with a 99% cache-hit discount) and $0.87 per 1M output tokens. This makes MiMo-V2.6-Pro one of the most cost-efficient models to deploy. MiMo-V2.6-Pro is an MoE model with 1.02T total parameters and 42B active parameters. Stay tuned for additional analysis of the model. Check out MiMo-V2.6-Pro full benchmarking breakdown here: https://artificialanalysis.ai

9/15Tue
  1. Cline · Blog62

    Cline releases the open-source desktop app Cline Desktop, aimed at open-weight models.

    Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.

    Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.

9/3Thu
  1. Cline · Blog74

    Cline 如何把 1100 万用户迁移到最大一次 harness 升级

    Cline 把 VS Code 扩展从约 76,000 行单体核心迁移到 Cline SDK,并自建灰度发布机制:一个安装包内打包 loader、legacy 和 next 两套扩展,由 PostHog 功能开关按百分比决定激活哪套,崩溃时自动回退到 legacy,开关可随时降到 0% 作为 kill switch。

    Awaiting translation

    Why it matters: Cline 官方复盘如何把 1100 万用户的 VS Code 扩展迁到新 harness,含灰度机制与前后指标对比。

8/27Thu
  1. Cline · Blog71

    Cline 实测八个模型做 IMO 2026:DeepSeek V4 Flash 以 0.12 美元拿到金牌线

    Cline 让八个模型在自家 harness 里做 IMO 2026 六道题,证明由 GPT-5.5 和 Claude Opus 5 双盲按 0–7 分制评分、Gemini 3.1 Pro 仲裁,金牌线为 29 分。

    Awaiting translation

    Why it matters: Cline 用同一套 harness 盲评八个模型做 IMO 2026,给出分数与单次成本对照,可看开源权重模型的实际性价比。

8/19Wed
  1. Cline · Blog71

    Cline's Evaluation Methodology and Trace for Open-Weight Models

    Cline has open-sourced its evaluation methodology for open-weight models, along with a hill-climbing score and trace worth over one thousand dollars, available for download and analysis. The post lays out five heuristics from the Hill Climber's Checklist: set a North Star metric, quantify noise, break down failure modes by task/model/vendor, don't assume more thinking is always better, and keep a private evaluation set.

    Why it matters: Cline shares its evaluation methodology and a trace worth over a thousand dollars, offering five transferable hill-climbing heuristics that teams building their own harness can reference.

8/11Tue
6/7Sun
  1. Thorsten Ball · Register Spill12

    Joy & Curiosity #89:Anthropic《When AI builds itself》、Ted Chiang 论 AI 意识与 Ladybird 停收公开 PR

    Thorsten Ball 的 Joy & Curiosity 通讯在写了三年后进入两三周、最多四周的暑期休更。本期链接包括 Anthropic 发布的《When AI builds itself》文档、Ted Chiang 关于 AI 是否具有意识的文章,以及 Ladybird 浏览器宣布不再接受公开 pull request,理由是 AI 工具已迅速改变开源信任的经济学。

    Awaiting translation