Skip to content

All updates

0 items today
9/23Wed
9/22Tue
  1. Sebastian Raschka72

    小米发布开源权重模型 MiMo-V2.6-Pro,在 Artificial Analysis 智能指数上以 46 分成为开源权重模型第一,每任务成本 0.13 美元,输入 0.435 美元/1M tokens、输出 0.87 美元/1M tokens,采用 1.02T 总参数、42B 激活参数的 MoE 架构。

    Awaiting translation

    QuotedArtificial Analysis@ArtificialAnlys

    MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46). At $0.13 per Intelligence Index task, it lands on the Intelligence vs. Cost per Task Pareto frontier @Xiaomi has just released MiMo-V2.6-Pro, an open weights model with major advances in intelligence over its predecessor, MiMo-V2.5-Pro (Intelligence Index: 26). Despite the improvement, it retains the same attractive pricing at $0.435 per 1M input tokens (with a 99% cache-hit discount) and $0.87 per 1M output tokens. This makes MiMo-V2.6-Pro one of the most cost-efficient models to deploy. MiMo-V2.6-Pro is an MoE model with 1.02T total parameters and 42B active parameters. Stay tuned for additional analysis of the model. Check out MiMo-V2.6-Pro full benchmarking breakdown here: https://artificialanalysis.ai

  2. Steve Yegge12

    前段时间,我经历了可能是我数学生涯中的高光时刻——我在一场数学争议中纠正了 Fable。Fable 5 当时在审阅我那篇关于 CI/CD 的博客文章,指责我在一个数学论断上玩弄修辞。在我纠正之后,我从没见过一个先进模型如此彻底地认怂。由于我就是那种所谓的“数学白痴”,我把这篇漫画献给每一位曾对我皱眉失望的数学老师。也就是说,献给所有数学老师。

    Awaiting translation

9/20Sun
9/19Sat
9/18Fri
9/16Wed
9/15Tue
  1. Lovable · Blog64

    Lovable open-sources OJ, a Rust preview engine that beats Vite on cold start and memory

    Lovable has released OJ, a preview engine written from scratch in Rust. It reads your existing vite.config.ts and runs real Vite plugins through a compatibility layer, all in a single binary, with no toolchain installed into the project.

    Why it matters: Lovable rewrote its preview engine OJ in Rust, sharing cold start and memory comparisons against Vite, plus canary data from production.

  2. Mitchell Hashimoto22

    Mitchell Hashimoto 演示了通过 CLI 控制 superlogical 多路复用器 Rex,GUI 能做的操作 CLI 都能做,且 CLI 功能更多,包括创建会话、运行命令、分屏、移动窗口、切换焦点、模拟键鼠输入、等待脚本等。Rex 还提供流式事件系统供任意客户端接入或脚本化,以及终端专属 API,可获取运行进程的 JSON 对象等信息。

    Awaiting translation

  3. Cline · Blog62

    Cline releases the open-source desktop app Cline Desktop, aimed at open-weight models.

    Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.

    Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.

9/14Mon
9/12Sat
9/11Fri
  1. GitHub Blog · Copilot30

    GitHub Copilot 应用新手指南:使用 diff、终端和浏览器面板

    GitHub Copilot 应用内置 diff、终端和浏览器三个面板,让开发者无需在编辑器、终端和浏览器之间切换即可完成 AI 编码闭环。diff 面板用绿色和红色高亮显示代码的新增、删除和修改,支持接受更改、留下评论或让 Copilot 继续修改;终端面板可直接运行项目命令并支持多窗口切换;浏览器面板则能预览界面并用 Pick & Polish 工具选中元素让智能体调整。

    Awaiting translation

  2. OpenAI Developer Blog · Codex66

    OpenAI on How to Rewrite Skills and Prompts for GPT-6 Astra

    In an official blog post, OpenAI lays out recommendations for adjusting Skills, AGENTS.md, and task prompts under GPT-6 Astra: Skill descriptions should be as short as possible and state clearly when they apply, and multi-flow Skills should use a root document for minimal routing instead of turning the Skill into an overly specific step-by-step checklist.

    Why it matters: OpenAI has published guidance on cleaning up Skills, AGENTS.md, and prompts under GPT-6 Astra, and it carries over to existing repository setups.

9/10Thu
  1. Cursor · Changelog76

    Cursor launches “Projects,” a feature that uses a coordinating agent to take on long-running development work

    Cursor introduces “Projects,” a feature built for long-running work like a single feature, a migration, or an entire application. It keeps context over months and delegates tasks to thousands of sub-agents. Projects are powered by cloud agents: the coordinating agent doesn’t write code, it only plans, assigns work, and hands back results, spinning up local agents when on-device testing is needed. Each project keeps a set of files synced between the cloud and local machines, steadily accumulating research findings, artifacts, and knowledge of the codebase.

    Why it matters: The official docs lay out the context-sharing and auto-triggering mechanisms for project-based multi-agent collaboration, which you can use to judge how long-running tasks get taken over.

9/9Wed
9/7Mon
9/6Sun
  1. Thorsten Ball · Register Spill15

    Thorsten Ball 谈 AI 写代码:是代码真的差,还是评审者的“天真干预”偏见

    Thorsten Ball 在 Register Spill 的 Joy & Curiosity #98 中借用《反脆弱》里的扁桃体切除研究,提出工程师对 Sol、Fable、Astra 等模型输出的“代码差、注释蠢”评价,可能源于“天真干预主义”偏见——AI 已在 20 分钟内端到端完成前后端改动、内外部文档和测试,并在无头浏览器中跑完全流程、附上录屏为证。

    Awaiting translation

9/5Sat
  1. GitHub Blog · Copilot71

    GitHub Copilot launches Project HydraFusion, using multi-model runtime orchestration to improve coding quality

    GitHub has launched Project HydraFusion as a research preview in the Copilot CLI. It uses runtime orchestration to pick an execution plan across models from multiple providers. Users select it just like any other model, and billing follows each model's standard rates.

    Why it matters: GitHub lays out three orchestration modes for HydraFusion and compares cost versus quality across three benchmarks, so you can judge the trade-offs of multi-model orchestration on real coding tasks.

9/4Fri
  1. OpenAI Developer Blog · Codex71

    How to Build a Game with Astra in Codex: From Void Explorer to Performance Tuning

    The author built the space exploration game Void Explorer in Codex with Astra, featuring 2,048 star systems and over 10,000 procedurally generated planets, and shared the full workflow from prompts to architecture, testing, and performance measurement.

    Why it matters: Using Astra in Codex, the author built an entire game and showed a transferable collaborative workflow that spans prompts, testing, and performance measurement.

  2. OpenAI Developer Blog · Codex28

    用 Codex 中的 Astra 做建筑可视化:从 Blender 场景到 Unreal Engine 5

    用户向 Codex 中的 Astra 描述一栋极简住宅的需求,Astra 通过 Blender Python API(bpy)生成可编辑 3D 场景,完成建筑、家具、材质、灯光与相机,并自行检查预览渲染、修正细节。项目从带家具的起居亭扩展为围绕庭院布局的 U 型单层住宅,含三间卧室、办公室、浴室和更衣室,随后导出到 Unreal Engine 5 探索实时漫游。

    Awaiting translation

9/3Thu
  1. Cline · Blog74

    Cline 如何把 1100 万用户迁移到最大一次 harness 升级

    Cline 把 VS Code 扩展从约 76,000 行单体核心迁移到 Cline SDK,并自建灰度发布机制:一个安装包内打包 loader、legacy 和 next 两套扩展,由 PostHog 功能开关按百分比决定激活哪套,崩溃时自动回退到 legacy,开关可随时降到 0% 作为 kill switch。

    Awaiting translation

    Why it matters: Cline 官方复盘如何把 1100 万用户的 VS Code 扩展迁到新 harness,含灰度机制与前后指标对比。

9/2Wed
  1. Cursor · Changelog66

    Cursor launches self-hosted machines, keeping tool execution within your own network

    Cursor supports self-hosted machines: code repositories, build artifacts, and secrets all stay on internal machines within your own infrastructure, and the agent handles tool calls locally. My Machines connects a single laptop or VM to a personal workflow, while Team Pools are named worker queues for teams or enterprises—scaling capacity up with requests and down when workers disconnect. Pools aren't tied to code repositories, and idle machines can sleep and then resume within a reconnection window.

    Why it matters: The official docs lay out pooled scheduling and sandbox integration for self-hosted machines, so readers can judge whether tool execution can stay within their own network.

9/1Tue
  1. Lovable · Blog38

    Lovable 接入 Fable 5.1:迭代修复最高提升 17%,成本降低 31%

    Lovable 现已接入 Fable 5.1,早期测试显示其在修复和改进现有应用上比 Fable 5 最高提升 17%,单任务成本最多降低 31%。该模型在中等和高推理强度下的 UI 与视觉设计质量最高提升 3.5%,并会在完成任务前打开浏览器运行应用进行自我验证。Lovable 正将卡住的会话以及更长、更复杂的任务路由到 Fable 5.1。

    Awaiting translation

8/30Sun
8/28Fri
8/27Thu
  1. Cursor · Changelog42

    Cursor 云端智能体无需代码仓库即可从零开始,支持保存到 Origin

    Cursor 云端智能体不再需要连接 GitHub 或其他第三方 SCM 提供商,用户可直接输入提示开始工作,Cursor 会在后台创建 Origin 代码仓库。满意后可点击“创建代码仓库”将工作保存到 Origin,并设置私有或内部可见性。Cursor 还能通过端口转发在浏览器中实时预览云端智能体环境,连接 Vercel 账户后点击“发布”即可生成可访问 URL。

    Awaiting translation

  2. Cline · Blog71

    Cline 实测八个模型做 IMO 2026:DeepSeek V4 Flash 以 0.12 美元拿到金牌线

    Cline 让八个模型在自家 harness 里做 IMO 2026 六道题,证明由 GPT-5.5 和 Claude Opus 5 双盲按 0–7 分制评分、Gemini 3.1 Pro 仲裁,金牌线为 29 分。

    Awaiting translation

    Why it matters: Cline 用同一套 harness 盲评八个模型做 IMO 2026,给出分数与单次成本对照,可看开源权重模型的实际性价比。

  3. GitHub Blog · Copilot30

    GitHub Copilot 应用入门:自动分诊 Dependabot pull request

    GitHub Copilot 应用支持用自动化分诊 Dependabot pull request:用自然语言描述任务,按风险分组、识别安全的补丁与次版本更新、核验 CI 状态并给出摘要。自动化可选手动、每小时、每天、每周或 issue 创建时触发,也可选择在云端或本地运行,每次运行记录都会保存。

    Awaiting translation

8/26Wed
  1. Cline · Blog71

    Building a Code Review Agent on the Cline Loop with the Cline SDK

    The Cline team built a code review agent with the Cline SDK, splitting review into two agent loops—review and judge—then using a driver script to batch-submit the surviving issues as a single COMMENT event to the GitHub PR.

    Why it matters: A full breakdown of the plugin, Hooks, and two-stage loop behind a code review agent, transferable to other automated review scenarios.