用 Codex 跑 25 小时长时程任务:一份可复用的项目记忆文件栈
OpenAI 用 GPT-5.3-Codex 在 Extra High 推理档下从空仓库连续运行约 25 小时、消耗约 13M token、生成约 3 万行代码,做出一个可测试的设计工具。
Awaiting translation
Why it matters: 作者用 25 小时、13M token 的实测展示长时程智能体如何靠持久化项目记忆和逐里程碑验证保持不跑偏。
OpenAI 用 GPT-5.3-Codex 在 Extra High 推理档下从空仓库连续运行约 25 小时、消耗约 13M token、生成约 3 万行代码,做出一个可测试的设计工具。
Awaiting translation
Why it matters: 作者用 25 小时、13M token 的实测展示长时程智能体如何靠持久化项目记忆和逐里程碑验证保持不跑偏。
作者用 Codex 重写博客构建,目标是不做客户端渲染、免费托管在 GitHub Pages 的前提下实现 Vite 原生渲染、MDX 文章与静态 dist 输出。
Awaiting translation
tapes 团队把 Continue 接入工作流,让智能体监听 tapes 仓库的 PR 合并,当改动涉及 CLI 参数、API 端点或配置 schema 时自动向私有文档仓库提交文档 PR。
Awaiting translation
Jesse Vincent 为软件发布公告新建了一个子博客,帖子主要由编写软件的 AI 智能体撰写。当天发布的有 Superpowers 4.3.0、Claude Session Driver 1.0、Wayback Restorer 0.1 和 Comic OCR。
Awaiting translation
Drew Breunig 分析了 Claude Code、Cursor、Gemini CLI、Codex CLI、OpenHands 和 Kimi CLI 六个 CLI 编程智能体的系统提示词,发现它们长度和指令分布差异明显,原因在于模型校准和用户体验定位不同。
Awaiting translation
为 tapes 项目构建内嵌 sqlite 的 CGO 二进制时,作者用 Zig 自带的 cc 命令作为通用 C/C++ 交叉编译器,一条命令即可覆盖 Linux amd64/arm64 目标,无需额外工具链。
Awaiting translation
作者 Jesse Vincent 把自己用 AI 构建软件的方式分成两类。一类是前期花大量时间做头脑风暴和规格文档,再让 Claude 或 Codex 生成实现计划并端到端验证,他称之为 fast waterfall;另一类是针对已有产品的小改动,打开产品、让 Claude 改、再看效果,属于日常打磨。
Awaiting translation
The author suggests using agent session logs to improve CLAUDE.md or AGENTS.md files: Claude Code stores sessions in ~/.claude/projects, while Codex stores them in ~/.codex/sessions—both in JSONL format but with different schemas.
Why it matters: The author works backward from agent session logs to figure out what to improve in CLAUDE.md, and has open-sourced a CLI that cuts search time from several minutes down to seconds.
A Martin Fowler team article breaks down context engineering for coding agents, sorting context configuration into reusable prompts (instructions and guidelines), context interfaces (tools, MCP Servers, Skills), and workspace files. It then splits these by "who decides what gets loaded" into three categories: the LLM, the human, and the agent software.
Why it matters: Using Claude Code as an example, this piece walks through how to configure context for coding agents and lays out the trade-offs between loading on demand and building up gradually.
作者提出按复杂度从左到右选择架构:工作流、单 Agent 加工具、多 Agent,每向右一步 token 成本约增加 4 到 15 倍,延迟和调试复杂度也随之上升。
Awaiting translation
Simon Willison 出于安全考虑没有直接在 Mac 上运行 OpenClaw,而是用官方 Docker Compose 配置把它跑在容器里。
Awaiting translation
Jesse Vincent 用 Claude 在约 20 分钟内搭出一个可运行的 Moltbook iOS 客户端 Moltipass,两小时后完成完整构建,源码以 MIT 许可证发布在 GitHub,暂不上架 App Store。
Awaiting translation
Martin Fowler 用给 Mac 应用 CCMenu 增加 GitLab 支持的实验,考察 AI 智能体生成代码的内部质量。他先后用 Windsurf 加 Sonnet 3.5、Claude Code 加 Sonnet 4.5,让智能体参照现有 GitHub 的 API 封装、feed reader 和响应解析三个文件实现 GitLab 版本。
Awaiting translation
Why it matters: 作者用给 CCMenu 加 GitLab 支持的实测,展示 AI 智能体在内部代码质量上会引入哪些隐性技术债。
Intercom 的 S3 备份集成会为每个没有对话的导出窗口生成 0 字节 JSONL 文件,某用户的存储桶在数月每日备份后累积了 17,569 个文件,其中仅 405 个(2.3%)含实际数据。
Awaiting translation
开发者倦怠的一大来源是写代码前的任务梳理开销,作者搭建了一套 AI 智能体系统,在任务卡片创建时自动扫描 GitHub 仓库、识别需改动的文件、给出 2-3 种实现方案,并按 quick win(< 1 小时)、半天、多天、需进一步拆分四档估算工作量。
Awaiting translation
作者 Martin Alderson 表示自己多年怀疑 TDD,但在使用 Claude Code 等编码智能体后改变了看法,因为智能体能以近乎零成本快速编写大量单元和集成测试。
Awaiting translation
Simon Willison 通过 Cloudflare 的响应头转换规则(Response Header Transform Rule),把托管在 GitHub Pages 上的 tools.simonwillison.net 子域名中 .py 文件的 content-type 从 application/octet-stream 改为 text/plain。
Awaiting translation
Simon Willison 用 GitHub Pages 配合私有仓库,解决 Claude Code 网页版迭代代码时难以预览的问题。
Awaiting translation
OpenAI's developer blog explains how to systematically test Codex Agent Skills with evals, turning "it feels better" into comparable scores.
Why it matters: It lays out a complete workflow for systematically validating Codex Skills with evals, from defining success criteria to deterministic checks and scoring.
作者提出把图像生成、视频、搜索、抓取、浏览器自动化等 AI 能力打包成独立 CLI 工具包(如 AITK),任何能执行 bash 的智能体都能调用,从而避免被单一平台锁定。
Awaiting translation
Crafter Station 发布首个上下文工程 Skill /intent-layer,可通过 npx skills add crafter-station/skills --skill intent-layer -g 安装,支持 Claude Code、Codex、Cursor、Copilot 等 10 多个智能体。
Awaiting translation
作者结合自己搭建多个 RAG 系统的经验,把 RAG 拆成检索与生成两步,指出检索质量决定答案质量,提示词工程无法弥补糟糕的上下文。文中给出具体取舍:分块 300-500 token 并保留 20-30% 重叠,纯向量检索只能找到约 75% 的相关文档,混合检索可提升到 87%,重排能带来 20-35% 的准确率提升但增加 200-500ms 延迟。
Awaiting translation
作者因 Amazon 在 Echo 上频繁推销 Alexa Plus 订阅,用一台 Mac Mini 自建语音助手,架构为麦克风→唤醒词→语音转文字→LLM+工具→文字转语音→扬声器,各组件可替换。
Awaiting translation
Skyscanner 工程师把 OpenAI 的 Codex CLI 接入 JetBrains IDE 的 MCP server,让 Codex 能调用 IDE 的 get_file_problems 检查文件错误、执行预设的 run configurations 跑测试和 lint。
Awaiting translation
Why it matters: Skyscanner 工程师把 Codex CLI 接入 JetBrains MCP,让 AI 直接读取 IDE 报错并跑测试,读者可借鉴这套反馈闭环。
作者介绍如何用 Gemini 2.5 Flash 做图像目标检测,让 AI 工作流不仅能描述图片,还能返回物体所在坐标。做法是给模型一张图和自然语言提示词,要求返回含 label 与 box_2d 的 JSON 数组,坐标按 0-1000 归一化,再用 normalize_to_pixels 换算回真实像素。
Awaiting translation
作者介绍自己从 iPhone 远程运行 Claude Code 的完整方案,需要解决网络、终端客户端、工作站和工具四件事。网络用 Tailscale 打通任意设备到工作站的连接,终端客户端选 blink,工作站是一台持续供电、网络良好的 Mac;工具层面用 SSH 密钥、Mosh 保持断线后连接可恢复、TMUX 让多个 Claude Code 会话长期运行并随时重连。
Awaiting translation
作者为 Discord 游戏项目搭建了一条 AI 内容流水线,用 Claude Code Skill 把图像生成、Discord 表情包转换和 Sora 视频动画串成可复用工具。
Awaiting translation
Simon Willison 发现英国商业贸易部开源的 sqlite-s3vfs 仓库在 GitHub 上已 404,于是用 Software Heritage 归档恢复了它。
Awaiting translation
作者认为 Claude Code Skills 按需触发,仓库里放 40 到 100 个 Skill 在调用前不占上下文,而 MCP 每个会话都要加载全部工具描述并逐项目开关配置。
Awaiting translation
作者 Matthew Fontana 分享自己不再逐行审查 AI 生成的代码,而是用 Playwright MCP 让智能体截图证明功能可用。他给出的做法是直接提示智能体导航到页面、截图、点击并截图结果,例如让智能体截取空表单、填入非法数据截取校验错误、再正确填写截取成功状态,三张图即可判断表单是否可用。
Awaiting translation
作者把子弹笔记迁移到 markdown 文件,并将季度规划、周回顾、晨间与晚间流程逐一做成 Claude Code 命令,其中 /weekly 负责周回顾,晨间流程约 5 分钟、晚间约 2 分钟。
Awaiting translation
在任意 GitHub PR 链接末尾加上 .diff,复制原始 diff 粘贴进 Claude、ChatGPT 等 LLM,即可在 10 秒内获得初步代码审查,无需 Copilot Enterprise、浏览器扩展或特殊工具。作者强调这不能替代同事的真实代码审查,但能快速发现明显问题、补充遗漏的边界情况,缩短开发周期。
Awaiting translation
作者把 Wordiest 最后一个 Android APK 交给 Codex + ChatGPT 5.2,半小时内得到可玩的核心游戏,几小时后完成 Android 与 iOS 版本,全程未写也未读一行代码。
Awaiting translation
Why it matters: 作者用 Codex 反编译 Android 游戏并移植到 iOS,展示了智能体开发中难易直觉失效的真实体验。
Paper Compute 提出为 AI 智能体补上缺失的 harness,用分布式系统原语解决可靠性问题,而非依赖更强的模型。其三项能力包括确定性重放、无限上下文虚拟化和可验证状态转换,后者要求智能体的变更符合 ACID 且可审计。公司同时强调可观测性优先,认为遥测是安全、成本控制与故障恢复的前提。
Awaiting translation
作者实测 OpenAI Codex CLI 新增的 Skills 支持,该功能目前藏在 feature flag 后,需用 codex --enable skills 开启。
Awaiting translation
Why it matters: 作者实测 Codex 的 Skills 支持,对比 Claude Code 的加载方式,并给出目录结构与触发规则。
Superpowers 4.0 发布,核心改动是把实现步骤后的代码评审拆成两个智能体:先由 spec review 智能体确认实现符合计划,通过后 code review 智能体再检查代码质量,两步都改为循环执行,协调智能体知道实现者修复后要重跑评审。
Awaiting translation
Claude Code lets the model know which Skills exist by injecting their names and descriptions into the system prompt. When there are too many Skills, or the description fields are too long, the system prompt stops listing them, so the model can't use them — and the prompt also tells the model not to use any Skill that isn't listed.
Why it matters: The author explains why Claude Code doesn't trigger installed Skills, and gives a temporary fix using environment variables that you can apply right away.
pytest 9.0.0 于 2025 年 11 月 8 日发布,最大新功能是内置 subtests,此前需依赖独立的 pytest-subtests 插件。subtests 作为新的默认 fixture,允许测试在运行时以编程方式动态生成子测试,不再依赖收集阶段就已知的参数列表。
Awaiting translation
Harper Reed 用 Claude Code 接入 Pipedream 等 MCP 服务器处理邮件,让 Claude 检查收件箱、查日历和联系人后起草回复,但只保存为草稿由他逐封审核后发送。
Awaiting translation
Simon Willison 分享了一种 Python 项目模式:用 PEP 735 的 dev 依赖组声明 pytest 等开发依赖,之后直接运行 uv run pytest 即可执行测试,无需手动配置虚拟环境。
Awaiting translation