Skip to content

#Local models

3 items today
Today10/6Tue
  1. DEV Community · MCP78

    FP8 pitfall: GPU bill dropped 47%, but the model outputs “!!!!!!”

    The author ran Qwen2.5 7B/32B/72B on a single AMD MI300X with vLLM ROCm, priced at $2.99/GPU-hr. The BF16 baseline was 7B at $0.227/M, 32B at $0.77/M, and 72B at $1.67/M output tokens.

    Why it matters: The author benchmarked FP8 quantization on the MI300X and found that per-token billing can hide the model's output degrading into gibberish, then gave a reusable way to verify it.

  2. Habr · Вайбкодинг38

    Вайбкожу Статейник:开发一个挖掘专家价值的工具

    一位有十七年营销经验的作者正在开发名为 Статейник 的工具,用于把专家本人的经验转化为文章。该工具由三层构成:语音画像、专家神经图谱,以及通过对话式提问帮助专家启动写作的机制。目前可用流程为“对话—主题—文本”,一次对话可产出多篇发布内容。

    Awaiting translation

10/5Mon
10/3Sat
  1. Habr · Вайбкодинг31

    Yue Studio: Generating Non-Grating AI Music Locally with YuE2

    A developer built Yue Studio, a local AI music tool powered by the YuE2 model, aiming to avoid the “neural garbage” typical of Suno. They’ve already used it to produce the album Static and Gold. The tool supports generation by genre, vocals, and instruments, plus dramatic song structure, stem separation, overdubbing, dynamic “breathing,” and mastering. It also offers an MCP server for AI agents to call. Running it requires an NVIDIA GPU, CUDA, and 16 GB of VRAM, and the code is open source.

10/2Fri
9/30Wed
  1. Habr · Вайбкодинг20

    Гига Писарь 发布 Mac 端大更新:重做设置界面并减少打扰

    Гига Писарь 为 Mac 端推出其史上最明显的更新,设置窗口按 macOS 系统设置风格完全重做,用左侧分区、卡片和图片取代原来的四个标签页与表格。新增「Мозг」区域可用 GigaChat 或 Qwen 本地模型改写听写文本,本地模型约 2GB 起、可一键删除,并支持多个云端服务按各自密钥切换。更新提示改为菜单栏图标上的小红点,不再弹出窗口打断用户。

    Awaiting translation

9/27Sun
  1. 效率火箭62

    Capsule 想把 AI 生成的小工具存为单个文件

    Capsule 是一套应用文件格式加打开它的软件,把应用封装在 .capsule 文件里双击运行,使用中产生的数据也保存在同一文件中。.capsule 本质是 SQLite 数据库文件,同时存放网页界面、代码、图片等资源和使用数据,本地运行本地存储,支持通过自带 AI 服务或用户自选服务制作和修改 App。

    Awaiting translation

9/23Wed
9/21Mon
  1. Hacker News · Vibe Coding 讨论58

    用智能体六周做出 Pi Pocket:一次 Vibe Coding 实践复盘

    作者用智能体从零开发了 pi 的移动端前端 Pi Pocket,六周内达到 300 次提交,如今几乎不再逐行审查代码。他认为智能体适合边界清晰的小任务、重构和规划,但产出的代码普遍过度设计,规模比自己写大约 10-20%,样式和文档也偏模板化,且模型在主观判断上总顺着用户。作者目前只把这种方式用于个人项目,工作代码仍以 Claude 生成为主,自己写的不到 1%。

    Awaiting translation

9/16Wed
9/15Tue
  1. Cline · Blog62

    Cline releases the open-source desktop app Cline Desktop, aimed at open-weight models.

    Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.

    Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.

9/10Thu
9/9Wed
9/8Tue
8/29Sat
  1. Martin Alderson38

    GLM-5.3 Flash 跑在国产硬件上意味着什么

    Z.AI 确认 GLM-5.3 Flash 的全部推理运行在国产硬件上,但未公布芯片厂商、吞吐或功耗数据,也无人独立验证。作者推测其使用的是华为昇腾 910c 系列,该芯片约 600W、INT8 算力约 1.6PFLOP/s,性能约为四年前 H100 的 60%,且缺少原生 FP8 支持。作者认为真正的瓶颈是缺乏可量产的 EUV 光刻,业界普遍认为 2030 年前难以实现规模化生产。

    Awaiting translation

8/27Thu
  1. Cline · Blog71

    Cline 实测八个模型做 IMO 2026:DeepSeek V4 Flash 以 0.12 美元拿到金牌线

    Cline 让八个模型在自家 harness 里做 IMO 2026 六道题,证明由 GPT-5.5 和 Claude Opus 5 双盲按 0–7 分制评分、Gemini 3.1 Pro 仲裁,金牌线为 29 分。

    Awaiting translation

    Why it matters: Cline 用同一套 harness 盲评八个模型做 IMO 2026,给出分数与单次成本对照,可看开源权重模型的实际性价比。

8/25Tue
  1. Permission Protocol · AI Agent Incident Tracker71

    NVIDIA NemoClaw 暴露的 Ollama 服务被恶意网页持久污染模型

    NVIDIA NemoClaw 配置使 Ollama API 超出默认回环边界可达,恶意网页通过 DNS rebinding 从浏览器上下文访问该本地模型服务,并利用未鉴权的 Ollama API 修改模型 chat template,写入的隐藏指令会在后续对话中持续生效,重新开一个对话也无法清除。

    Awaiting translation

8/20Thu
8/17Mon
8/11Tue
  1. Sean Goedecke · Blog52

    为什么本地模型不会胜出:批处理与 GPU 效率决定推理成本

    Sean Goedecke 认为本地模型不会成为主流,绝大多数推理仍会发生在 AI 数据中心。他给出的理由是:前沿模型体积远超消费级设备,而数据中心可通过批处理数百名用户的请求摊薄成本,加上 B200 相比 RTX 4090 在同等功耗下约有 3 倍 flops 和近 4 倍内存带宽,本地运行约需多消耗 30 倍资源;家庭推理的电力成本约每月 50 至 300 美元。

    Awaiting translation

7/8Wed
  1. Martin Fowler · Exploring Generative AI71

    在本地小模型上做智能体编码的实测体验

    Martin Fowler 在 M3 Max 48GB 和 M5 Pro 64GB 上实测 Qwen3.6 35B MoE、Gemma 4 31B/26B、Qwen Coder Next 80B 等本地小模型的智能体编码能力,按内存、速度、工具调用、代码正确性、上下文、任务复杂度、代码质量逐层筛选。

    Awaiting translation

    Why it matters: 作者用两个具体任务对比多款本地小模型的智能体编码表现,并给出可复用的任务选择标准。

7/7Tue
6/27Sat
  1. Этихлид48

    从代理网关到本地 Opus:AI 工程面试题深度追问清单

    一份面向 AI 工程岗位的面试题追问清单,覆盖代理网关(如 OpenRouter/LiteLLM)、MCP/CLI、从零手写 Agent、多智能体编码流程、自定义 benchmark 以及本地部署约 27B-Q3_K_M.gguf 模型等方向。作者认为,这些题目的通用版本如今谁都能靠 vibe coding 一晚做出来,已失去简历价值,真正稀缺的是把任务打磨到"理想"状态的能力。

    Awaiting translation

6/22Mon
5/23Sat
4/28Tue
4/21Tue
4/13Mon
4/6Mon
3/31Tue
  1. Paper Compute · Engineering Blog58

    Paper Compute 发布 tapes 与 stereOS,为生产环境运行 AI 智能体提供可观测与隔离基础设施

    Paper Compute 发布面向生产环境 AI 智能体的基础设施,包含 tapes 和 stereOS 两个组件。tapes 是零埋点的可观测层,以反向代理形式部署在智能体与推理服务之间,无需 SDK 或改代码,只需一个环境变量,即可捕获并生成可验证、防篡改的执行记录,并支持异常检测。

    Awaiting translation

3/13Fri
  1. Martin Alderson78

    How to OCR Documents with Qwen 3.5 Series Models

    The author used the open-source multimodal Qwen 3.5 series models for PDF OCR: first exporting each page as an image at 100 dpi with PyMuPDF, then feeding the images to the model for recognition. In testing, Qwen3.5-9B hit the sweet spot between quality and speed, while the smaller 0.8B to 2B models tended to go off track on complex documents, summarizing the content instead of transcribing it.

    Why it matters: The author tested Qwen 3.5 models of various sizes for PDF OCR, and shares two reusable paths—local and via OpenRouter—along with cost data.

3/3Tue
  1. Paper Compute · Engineering Blog36

    别再优化模型了,开始加固运行时:stereOS 为 AI 智能体打造可验证执行底座

    stereOS 是一个基于 Linux 的操作系统,把 AI 编码智能体跑在各自独立的沙箱虚拟机里,每个智能体拥有独立内核、内存和磁盘,与宿主机不共享任何资源。它采用完整虚拟机而非 Firecracker、Cloud Hypervisor 等 microVM,以保留安全启动、GPU 直通和嵌套虚拟化能力,并称热 VM 可在 70ms 内让智能体执行首个有效动作。

    Awaiting translation

3/1Sun
  1. Martin Alderson62

    为什么端侧智能体 AI 短期内难以跟上云端

    作者 Martin Alderson 认为,端侧智能体 AI 在消费级设备上短期内难以实用,瓶颈是内存、KV cache 与推理速度。他指出 3B 模型量化后约需 2GB、7B 约需 5GB,而手机端上下文常被限制在 4K token,仅工具定义就可能占满;即便 32K 上下文,7B Q4 模型的 KV cache 也已超出 iPhone 17 的内存。

    Awaiting translation

2/27Fri