Skip to content

#Agent

0 items today
10/5Mon
10/4Sun
  1. DEV Community · Claude Code66

    Claude 桌面端定时任务卡在审批三天未执行,作者改用终端 cron 恢复晨间更新

    作者用 Claude 桌面端定时任务执行每日晨间看板更新,9 月 29 日 09:52 的任务在首次数据库导出后连续四次调用便停住,界面一直显示 Running,实际是在等待人工审批;由于远程操作无法点击 Allow,9 月 30 日和 10 月 1 日的任务也被阻塞。

    Awaiting translation

  2. DEV Community · Claude Code85

    How to Stop an AI Coding Agent from Declaring a Task Done Too Early

    The author runs a fully autonomous implementation system where an orchestrator hands out tasks to parallel implementation agents (built on Claude Code). At first, agents could mark their own tasks as complete, which led to problems like tests never being run, acceptance criteria not being met, assertions loosened to make tests pass, and hardcoded return values.

    Why it matters: The author solved the problem of agents declaring completion too early with a three-layer design: checkable acceptance criteria, completion reports backed by evidence, and read-only validation agents.

  3. 宝玉82

    Drawing on a podcast episode, Baoyu walks through how Lauren Tan, who works on Grok Bot at SpaceXAI, merged 2500 PRs in a single month: at night she lets the AI check and merge on its own, then spot-checks the next morning instead of reviewing each one.

    Quotedlauren@poteto

    i had a lot of fun chatting with @mattpocockuk today about how i was able to land 2,500 PRs last month! Matt is a wonderful interviewer so i think the interview turned out really interesting both of our skill plugins work great together, so i recommend giving both a try and picking the best skills that suit your workflow https://www.youtube.com/watch?v=MN9dGgmLyso

    Why it matters: Using Lauren Tan's practice of merging 2500 PRs in a month, Baoyu explains that skipping individual reviews rests on validation Skills and rule constraints, and lays out the conditions under which he'd apply the same approach.

  4. 佬刘AI78

    Planning with GPT, Execution with DeepSeek: A Cost-Saving Two-Model Workflow

    The author has GPT (gpt-6.1-sol, reasoning tier high) handle planning, key decisions, and acceptance in Codex, and calls DeepSeek-V4.1-Flash through the official DeepSeek Harness to write code, run experiments, and fix bugs—together producing a local PDF toolkit with five working features.

    Why it matters: The author splits the work between GPT for planning and DeepSeek for execution to get a PDF toolkit running, and shares three prompts plus cache usage data that can carry over to cutting costs on long tasks.

  5. Simon Willison · Coding Agents39

    我们需要给几乎所有服务加上默认硬性预算上限

    Simon Willison 呼吁按用量付费的 API 和服务默认提供硬性预算上限,即达到 $X/月后直接切断并返回错误,而非只发警告邮件。他指出 AWS 已在 9 月 16 日推出月度支出上限功能,项目达到限额后当月暂停,Google Cloud 也在 7 月上线了 Spend Caps。他认为硬性上限应成为默认选项,取消需用户主动勾选确认。

    Awaiting translation

  6. DEV Community · Vibe Coding74

    Android 上 Vibe Coding 的问题:AI 生成代码的幻觉、协程泄漏与安全数据

    作者梳理 AI 生成 Android 代码的常见问题,并引用多项研究数据:USENIX Security 2025 分析 223 万个生成代码样本、16 个模型,开源模型包名幻觉率平均 21.7%,商业模型 5.2%;CodeRabbit 分析 470 个开源 PR 发现 AI 代码缺陷率是人类代码的 1.7 倍,性能问题接近 8 倍。

    Awaiting translation

10/3Sat
  1. TonyBai65

    Agent 引擎 Pi 发布 1.0,把 MCP 收进核心并推出 Pi Durable

    Agent 引擎 Pi 于 10 月 1 日发布 1.0 版本,GitHub 星标已达 11.1 万。此前宣称不支持 MCP 的团队这次把 MCP 收进核心,并推出 Codemode,让模型写脚本编排工具,实测 331 次调用过程不占上下文;同时支持虚拟模型,例如 Claude Opus 负责规划、Jev 负责决策、GPT 负责写代码,切换自动并统计花费。

    Awaiting translation

  2. TonyBai80

    Pi 1.0 is out: the Agent engine behind OpenClaw now takes in MCP and ships Pi Durable with crash recovery

    On October 1, the Earendil team released the Agent Harness Pi 1.0, with roughly 11.1 stars and 1.4 forks on GitHub, plus an experimental new package called Pi Durable.

    Why it matters: Pi 1.0 and Pi Durable bring distributed concepts like checkpoints, idempotent commits, and ownership trees into the Agent runtime, which you can use to weigh the engineering trade-offs of long-running Agents.

  3. 宝玉62

    剑桥大学 AI 科学与政策项目(CASP)9 月发布一篇论文,Hinton、Bengio 等 20 多位研究者联名讨论 AI 研发被自动化后是否会引发智能爆炸,Hinton 在 X 上推荐了该论文。

    Awaiting translation

    QuotedGeoffrey Hinton@geoffreyhinton

    The idea of an intelligence explosion caused by recursive self improvement has been around for a long time but until very recently it did not seem imminent. Now many leading researchers think it may happen quite soon. You can read our paper about it here: https://casp.ac/reports/intelligence-explosion

  4. Cursor Forum · Showcase60

    Pinpoint:点击页面元素并说明改动,Cursor 即可获得上下文

    作者因厌倦粘贴截图和用文字描述“左边第三个按钮”,开发了 Pinpoint,一个本地 MCP server 加 Chrome 扩展。在开发站点点击元素并输入要改的内容后,Cursor 会收到评论以及选择器、DOM 路径、计算样式、React/Vue 组件链、源文件提示和该元素的裁剪截图。

    Awaiting translation

10/2Fri