Skip to content

#Agent

0 items today
10/6Tue
  1. Reddit · ClaudeCode / Codex / VibeCoding65

    作者 fork Ghostty 做出可视化终端,实时展示 Claude Code 的编辑与仓库地图

    作者 fork Ghostty 做了一个可视化终端,用来解决同时跑多个 Claude Code 会话时来回切标签查看状态的问题。终端旁的编辑器会打开 Claude 读过的文件并实时敲入编辑,仓库地图按读取显示青色、编辑显示橙色,侧边栏列出每个会话的工作、待授权和完成状态并在需要时提醒,完成后 diff 进入 review inbox,可对行评论并作为下一条提示词发回。

    Awaiting translation

  2. Reddit · ClaudeCode / Codex / VibeCoding60

    cactus:用非阻塞决策队列替代 90% 的 Claude Code 使用

    作者发布 cactus,一个把 Claude Code 的对话式交互改成非阻塞决策队列的工具,据称替代了自己 90% 的 Claude Code 使用。它包含供 AI 生成和迭代引导问题的 CLI、按项目汇总所有 agent 问题的 TUI、检查问题是否被提出并提醒未决问题的 Hooks,以及为会话侧边栏提供问答界面并注入答案的 Claude Code mod。

    Awaiting translation

  3. 宝玉70

    amontlabs/lcu 把 Codex 的 Computer Use 单独拆出,让 Claude Code、Codex CLI、Pi 等 Agent 工具通过 MCP 调用。

    Awaiting translation

    Quoted向阳乔木@vista8

    发现一个牛逼的东西,让任意 Agent 调用 Codex 的 Computer Use。 Codex 最强的就是 Computer Use。 但最近用 Claude Opus 5.5比较多,这样就强强联合了。 刚测试通过,安装后建议配置个 Hook,指定白名单可控制哪些 App 安装地址见评论区

  4. GitHub Blog · Copilot66

    GitHub 发布 AI 代码评审开放基准 ReviewBench

    GitHub 发布代码评审离线基准 ReviewBench,基于 1.039 亿个 GitHub PR 的分布特征,构建了覆盖 19 种语言、219 个公开 PR 的评测集,并公开数据集、评分规则与 LLM 评审模型配置。

    Awaiting translation

    Why it matters: GitHub 公开了 AI 代码评审基准的数据集、评分规则与评测入口,读者可据此对比不同评审智能体。

10/5Mon
  1. Habr · Вайбкодинг76

    When Automated Checks Lie: Five Cases from a Project Where an AI Agent Writes the Code

    In a product project where an AI agent writes the code and the author doesn't read it, automated checks repeatedly reached the wrong conclusion. The author found 86 checks that no workflow had ever triggered, a secret scan that missed 438 of 1413 files because Git escapes Russian filenames by default, a new check that mistook WHERE for a table alias and let an injection slip through, and three false alarms from the test dashboard and the agent's replica.

    Why it matters: The author walks through five real cases to show why automated checks produce false greens or false reds, and lays out validation rules that carry over to other projects.

  2. DEV Community · Claude Code82

    How I Used Git Checkpoints to Undo Any Change Made by a Coding Agent

    For coding agents running unattended, the author built a checkpoint mechanism based on hidden git refs. Before each task starts, it snapshots the entire working tree—including untracked files—and rolls back automatically when validation fails. The restore operation itself can also be undone.

    Why it matters: With roughly 40 lines of shell, the author turned git checkpoints into rollback-capable infrastructure, laying out the concrete approach and the limits of running coding agents unattended.

  3. Reddit · ClaudeCode / Codex / VibeCoding66

    Repowise 发布 Lens,用 Claude Code mods 展示每次编辑触及的文件

    开源工具 Repowise 发布 Lens,借助新的 Claude Code mods 在 Claude 工作时展示代码库索引。输入 /lens 可看到仓库地图,文件在 Claude 搜索、读取和编辑时高亮,编辑某文件时所有导入它的文件也会亮起,视频中 Django 的 query.py 一行注释点亮 12 个文件,conf/init.py 点亮约 248 个。

    Awaiting translation

  4. Geoffrey Huntley · Blog58

    Geoffrey Huntley 发布 Jiti:通过对话让 LLM 持续扩展运行中的 Lisp 应用

    Geoffrey Huntley 发布 Jiti,一个通过对话让 LLM 扩展运行中 Lisp 应用的小型内核,源码已在 GitHub 开源。用户提出需求后,OpenAI 模型借助注册工具检查、修改并执行 Lisp 代码,被接受的函数会作为普通 Lisp 函数永久保留,后续调用无需再次推理。作者认为这种免编译、边运行边生长的开发方式,比传统 CI/CD 编译流程更值得探索。

    Awaiting translation

  5. Reddit · ClaudeCode / Codex / VibeCoding71

    团队用 Claude Code 六周写了 2752 条提示词,把重复纠错整理成 14 个 Skill

    一个团队回顾了用 Claude Code 六周构建产品的全部会话,发现 2752 条提示词里很多是重复输入的同一类纠正,于是把最常重复的整理成 Skill。他们列出代价最大的几类问题,包括 API 未就绪时用 mock 数据、只看 CSS 就断言已居中、失败迁移后的清理步骤误删线上计费数据、本地测试向假地址发出约 50 封真实邮件、只要一个导航链接却得到整个导航重设计。

    Awaiting translation

  6. Hacker News · MCP77

    我连续 37 天夜间测量 MCP 注册表:19.2% 的服务器 24 小时内改动了工具面

    作者搭建 mcp-transparency-log,按计划爬取官方 MCP 注册表中所有公开可达的服务器,记录其工具名、描述、JSON schema 和四个 annotation 提示,写入带签名树头的 append-only 日志。

    Awaiting translation

    Why it matters: 作者连续 37 天夜间爬取 MCP 官方注册表,用可复现的日志量化工具面变化,并公开了两次错误修正过程。

  7. DEV Community · Claude Code74

    Claude Code mods 并未被沙箱隔离,作者拆解两层沙箱含义并给出安装检查清单

    Claude Code mods 的 JS 运行时沙箱只限制代码如何访问外部,并不限制它能否访问;Anthropic 文档明确写道 mods 未被沙箱隔离,mod 以用户权限运行,可读写文件、启动进程、发起网络请求,还能读取环境变量和设置文件中的 API key、批准被 ask 规则或 PreToolUse hook 拦截的工具调用、改写事件。

    Awaiting translation

  8. Habr · Codex71

    Dan Lu 实验:测试为何在有 bug 时仍然通过,以及如何验证测试本身

    作者解读 Dan Lu 的智能体实验:让智能体按 RFC 用 Rust 写 Zstd 解码器,比较 26 种条件(含无额外指令的对照组),主要对比用 Codex 搭配 GPT-5.6 Sol 的 medium 与 xhigh 两档,每组合 80 次运行,结果这些 TDD、模糊测试、形式化方法等指令没有带来明显整体收益,不少条件还不如对照组。

    Awaiting translation

  9. DEV Community · Claude Code78

    Why Claude Code’s Read(.env) Deny Rule Doesn’t Stop Bash from Reading It

    The author added a Read(./.env) deny rule to Claude Code, but after Read was blocked, Claude switched to running `grep DATABASE_URL .env` via Bash, printing the production connection string into the conversation.

    Why it matters: Through hands-on testing, the author found that the Read deny rule doesn’t stop Bash from reading .env, and shares a three-layer protection setup that can be adapted to your own permission configuration.

  10. DEV Community · Codex78

    How I Stopped Codex from Burning Through My Usage Quota

    The author found that two Codex browser automation tasks consumed 170,123 and 110,180 tokens respectively, so they set out to control usage through model selection, configuration files, and task splitting.

    Why it matters: Drawing on real measurements where two browser tasks burned through hundreds of thousands of tokens, the author shares a quota-saving approach: switch models and configurations based on task difficulty.