Skip to content

Open and local models

Coding with open-weight and local models: integration, performance comparisons, and deployment experience.

Latest curated items

Items 1–10 · 10 total
Today10/6Tue
  1. DEV Community · MCP78

    FP8 pitfall: GPU bill dropped 47%, but the model outputs “!!!!!!”

    The author ran Qwen2.5 7B/32B/72B on a single AMD MI300X with vLLM ROCm, priced at $2.99/GPU-hr. The BF16 baseline was 7B at $0.227/M, 32B at $0.77/M, and 72B at $1.67/M output tokens.

    Why it matters: The author benchmarked FP8 quantization on the MI300X and found that per-token billing can hide the model's output degrading into gibberish, then gave a reusable way to verify it.

10/4Sun
  1. DEV Community · Cursor76

    Cursor ships Composer 2, and the API response strings give away its undisclosed Kimi K2.5 base

    On March 20, 2026, developer Fynn was debugging Cursor's OpenAI-compatible endpoint when the returned model ID came back as accounts/anysphere/models/kimi-k2p5-rl-0317-s515-fast — evidence that Composer 2 was post-trained with reinforcement learning on top of Moonshot AI's Kimi K2.5. The tweet hit 44.4 views within a day.

    Why it matters: One API debugging session ties together Cursor's undisclosed Kimi base, the licensing attribution dispute, and the cost landscape for Chinese versus U.S. models — a look at how the industry handles disclosure.

9/15Tue
  1. Cline · Blog62

    Cline releases the open-source desktop app Cline Desktop, aimed at open-weight models.

    Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.

    Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.

9/3Thu
  1. Cline · Blog74

    Cline 如何把 1100 万用户迁移到最大一次 harness 升级

    Cline 把 VS Code 扩展从约 76,000 行单体核心迁移到 Cline SDK,并自建灰度发布机制:一个安装包内打包 loader、legacy 和 next 两套扩展,由 PostHog 功能开关按百分比决定激活哪套,崩溃时自动回退到 legacy,开关可随时降到 0% 作为 kill switch。

    Awaiting translation

    Why it matters: Cline 官方复盘如何把 1100 万用户的 VS Code 扩展迁到新 harness,含灰度机制与前后指标对比。

8/27Thu
  1. Cline · Blog71

    Cline 实测八个模型做 IMO 2026:DeepSeek V4 Flash 以 0.12 美元拿到金牌线

    Cline 让八个模型在自家 harness 里做 IMO 2026 六道题,证明由 GPT-5.5 和 Claude Opus 5 双盲按 0–7 分制评分、Gemini 3.1 Pro 仲裁,金牌线为 29 分。

    Awaiting translation

    Why it matters: Cline 用同一套 harness 盲评八个模型做 IMO 2026,给出分数与单次成本对照,可看开源权重模型的实际性价比。

8/19Wed
  1. Cline · Blog71

    Cline's Evaluation Methodology and Trace for Open-Weight Models

    Cline has open-sourced its evaluation methodology for open-weight models, along with a hill-climbing score and trace worth over one thousand dollars, available for download and analysis. The post lays out five heuristics from the Hill Climber's Checklist: set a North Star metric, quantify noise, break down failure modes by task/model/vendor, don't assume more thinking is always better, and keep a private evaluation set.

    Why it matters: Cline shares its evaluation methodology and a trace worth over a thousand dollars, offering five transferable hill-climbing heuristics that teams building their own harness can reference.

8/11Tue
7/8Wed
  1. Martin Fowler · Exploring Generative AI71

    在本地小模型上做智能体编码的实测体验

    Martin Fowler 在 M3 Max 48GB 和 M5 Pro 64GB 上实测 Qwen3.6 35B MoE、Gemma 4 31B/26B、Qwen Coder Next 80B 等本地小模型的智能体编码能力,按内存、速度、工具调用、代码正确性、上下文、任务复杂度、代码质量逐层筛选。

    Awaiting translation

    Why it matters: 作者用两个具体任务对比多款本地小模型的智能体编码表现,并给出可复用的任务选择标准。

5/7Thu
  1. Permission Protocol · AI Agent Incident Tracker80

    仿冒 OpenAI 仓库在 Hugging Face 登顶热门榜并获 24.4 万次下载后投递窃密木马

    Hugging Face 上一个仿冒 OpenAI Privacy Filter 的仓库登上热门榜第一、获得 244,000 次下载,随后在安装该模型的 Windows 机器上执行窃取凭据的 infostealer。

    Awaiting translation

    Why it matters: 复盘 Hugging Face 上仿冒 OpenAI 仓库的投毒链条,展示热门榜如何被当作信任信号利用。

3/13Fri
  1. Martin Alderson78

    How to OCR Documents with Qwen 3.5 Series Models

    The author used the open-source multimodal Qwen 3.5 series models for PDF OCR: first exporting each page as an image at 100 dpi with PyMuPDF, then feeding the images to the model for recognition. In testing, Qwen3.5-9B hit the sweet spot between quality and speed, while the smaller 0.8B to 2B models tended to go off track on complex documents, summarizing the content instead of transcribing it.

    Why it matters: The author tested Qwen 3.5 models of various sizes for PDF OCR, and shares two reusable paths—local and via OpenRouter—along with cost data.