Skip to content

Practice and tutorials

Best practices, reproducible workflows, and tutorials with commands and configuration examples.

Latest curated items

Items 81–100 · 117 total
4/24Fri
  1. Jesse Vincent66

    Prime Radiant 发布 Greenfield 与 Iterative Development 研究预览

    Prime Radiant 发布两款新技术的研究预览:Greenfield 把现有软件(代码库、文档、API 客户端等)转成行为规格语料,Iterative Development 则是一套基于 Superpowers 的智能体方法论,把大规格拆成需求、打包成开发 epic 后交给编码智能体实现。

    Awaiting translation

    Why it matters: Prime Radiant 公开两套工具的研究预览,读者可了解从旧代码库提取行为规格再驱动智能体重建产品的思路。

4/22Wed
  1. Augment Code · Blog88

    Augment Code Tests AGENTS.md: A Good File Is Like a Model Upgrade, a Bad One Is Worse Than Nothing

    Augment Code pulled dozens of AGENTS.md files from its own monorepo and used its internal benchmark suite AuggieBench to compare how the same tasks performed with and without the file. The best files delivered a quality boost equivalent to upgrading from Haiku to Opus, while the worst made the output worse than having no AGENTS.md at all.

    Why it matters: Augment Code used internal benchmarks to quantify how much the different ways of writing AGENTS.md actually differ, so readers can adjust their own repo's documentation structure accordingly.

4/15Wed
  1. Kondasamy Jayaraman · Engineering Blog78

    12 Agent-Building Patterns Distilled from the Claude Code Source Leak

    After analyzing the architecture that surfaced in the Claude Code source leak, the author argues that its core isn't a secret algorithm but a while loop plus a tool dictionary in under 30 lines of Python, driven by stop_reason !

    Why it matters: From the leaked source, the author distills 12 composable agent-engineering patterns and lays out a four-week path to get started, useful for checking your own implementation for gaps.

4/7Tue
  1. Jesse Vincent74

    如何用门禁而非规则约束 AI 智能体行为

    Jesse Vincent 在构建 Superpowers 时提出,提示词中的门禁(gate)比规则(rule)更能约束 AI 智能体:规则留有自我说服的退出路径,门禁则要求满足条件才能进入下一步。

    Awaiting translation

    Why it matters: 作者用自家智能体的实例区分规则与门禁,给出可迁移的提示词写法,帮助减少智能体跳过验证的行为。

4/6Mon
  1. 宝玉78

    Claude Code Token-Saving Guide: Be Careful with the 1M Context—Neither Never Opening a New Session Nor Always Opening One Is Right

    Baoyu walks through Claude Code's prompt caching mechanism to explain why quotas burn so fast, and lays out rules for saving tokens. He points out that caching only applies to prefixes, the main agent's cache window is 1 hour, and sub-agents' is 5 minutes. Reading from cache costs about one-tenth of recomputing, so frequent /clear actually triggers a full-price context rebuild. The rule of thumb: if the cache is still warm and the task hasn't changed, keep chatting; only start a new session when the cache has expired, the task has shifted, or there's too much context noise.

    Why it matters: Starting from the prompt caching mechanism, this explains Claude Code's quota consumption and gives the criteria for deciding whether to continue a session or start over, plus configuration you can copy.

4/5Sun
  1. Drew Breunig78

    How Claude Code assembles system prompts

    Based on the Claude Code source code that leaked unexpectedly last week, Drew Breunig mapped out how the system prompt is assembled: components fall into two categories—always included and conditionally included—and shift based on toggles like output_style, repl_mode, user_type_ant, skills_enabled, and mcp_connected.

    Why it matters: The author breaks down the dynamic assembly logic behind Claude Code's system prompt, showing how conditional context engineering works in practice.

3/17Tue
  1. 宝玉87

    How Anthropic’s Team Uses Claude Code Skills: Nine Categories and Writing Tips

    Thariq Shihipar, an engineer on Anthropic’s Claude Code team, summed up what the team learned from using hundreds of active Skills internally, sorting them into nine categories: library and API references, product validation, data acquisition and analysis, business process automation, code scaffolding, code quality and review, CI/CD and deployment, operations runbooks, and infrastructure operations.

    Why it matters: Anthropic’s internal classification system for hundreds of Skills, along with its writing tips, can be adapted to help teams design their own Skills.

  2. Paper Compute · Engineering Blog78

    日志即自愈反馈回路:用遥测让智能体跨会话积累经验

    作者让智能体在 stereOS 虚拟机里用 PyBoy 无头运行宝可梦红,速度约为实时的 100 倍,智能体自己输出 NAV、BATTLE、BACKTRACK 等日志前缀,这些日志经 tapes 代理流入 Kafka,再由 Flink SQL 做 STUCK_LOOP、TOKEN_SPIKE 异常检测,JSONL 与 DuckDB 负责跨会话查询。

    Awaiting translation

    Why it matters: 作者用终端里跑宝可梦的智能体做实验,展示日志如何变成跨会话的观测记忆并反哺下一轮运行。

3/16Mon
  1. 宝玉78

    The 8 Levels of Agent Engineering: From Tab Completion to Autonomous Agent Teams

    Bassim Eledath breaks the practical path of AI-assisted programming into 8 levels, from tab completion and agentic IDEs to context engineering, compound engineering, MCP and Skills, Harness Engineering, background agents, and finally autonomous agent teams.

    Why it matters: The author lays out AI-assisted programming as 8 levels, from tab completion to autonomous agent teams, so readers can figure out where their own team stands.

3/13Fri
  1. Ryan Lopopolo71

    别再把代码当成最终产物:Symphony 作者谈规范驱动开发

    Ryan Lopopolo 以 Symphony 为例提出,代码不应被视为最终产物,规范才是分发物。Symphony 是一个基于 issue tracker 的智能体编排系统,先以 SPEC.md 形式分发,Elixir 参考实现只是衍生产物,README 邀请读者把 SPEC.md 交给自己的编码智能体,用任意语言重建。

    Awaiting translation

    Why it matters: 作者用 Symphony 的 spec 蒸馏循环说明代码只是可替换产物,为规范驱动开发提供了一条可迁移的路径。

  2. Ryan Lopopolo74

    智能体时代的生产函数变了:实现不再稀缺,验证才是

    作者认为,软件组织一直把人的实现时间当作稀缺投入,智能体打破了这个假设,因此策略必须可执行、验证必须随实现规模扩展。他以在大型 TypeScript 代码库开启 ESLint 的 no-await-in-loop 规则为例,发现 600 处违规,过去这需要昂贵的迁移,现在一个 PR 就能完成修复并补齐测试覆盖。

    Awaiting translation

    Why it matters: 作者以开启 ESLint 规则、迁移 600 处违规的亲身实践,说明智能体时代实现成本下降后,约束与验证为何成为新的稀缺环节。

  3. Martin Alderson78

    How to OCR Documents with Qwen 3.5 Series Models

    The author used the open-source multimodal Qwen 3.5 series models for PDF OCR: first exporting each page as an image at 100 dpi with PyMuPDF, then feeding the images to the model for recognition. In testing, Qwen3.5-9B hit the sweet spot between quality and speed, while the smaller 0.8B to 2B models tended to go off track on complex documents, summarizing the content instead of transcribing it.

    Why it matters: The author tested Qwen 3.5 models of various sizes for PDF OCR, and shares two reusable paths—local and via OpenRouter—along with cost data.

3/9Mon
  1. Paper Compute · Engineering Blog78

    我让 Agent 玩 1000 回合宝可梦,它始终没走出卧室

    作者构建的宝可梦红版自主 Agent 跑了 1000 回合仍停在卧室,原因是它把背景瓦片图地址 0xC4F2 当成文本框状态标志,一直按 A 键却不知道卡住。

    Awaiting translation

    Why it matters: 作者用 1000 回合卡在卧室的失败复盘,说明记录并回放 Agent 会话状态比改提示词更能定位静默故障。

  2. Jesse Vincent67

    Superpowers 5 发布:新增可视化头脑风暴与 spec 评审循环

    Superpowers 5 发布,作者称最喜欢的改动是 Visual Brainstorming 伴随工具,它会在智能体认为有内容需要展示时提示用户,通过本地 web 服务器加载智能体写出的 HTML 片段,并把浏览器里的点击和反馈回传给智能体,以替代 Claude 常生成的 ASCII 图。

    Awaiting translation

    Why it matters: 作者是 Superpowers 维护者,文中说明了 5.0 的视觉头脑风暴、spec 评审循环和子智能体开发三项变化,可据此判断是否值得接入现有工作流。

  3. OpenAI Developer Blog · Codex87

    OpenAI 如何用 Skills 加速 Agents SDK 仓库维护

    OpenAI 用 Codex 配合仓库内的 Skills、AGENTS.md 和 GitHub Actions 维护 Agents SDK 仓库,把验证、发布准备、示例集成测试和 PR 评审变成可重复流程。

    Awaiting translation

    Why it matters: OpenAI 官方公开了用 Skills、AGENTS.md 和 GitHub Action 维护 Agents SDK 仓库的完整配置,可迁移到其他开源项目。

3/5Thu
  1. Lovable · Blog71

    Lovable 如何每分钟路由十亿 token:多回退链与项目级粘性负载均衡

    Lovable 的基础设施团队公开了其 LLM 供应商负载均衡方案,用于在峰值每分钟超过十亿 token 的流量下避免“model provider unavailable”。

    Awaiting translation

    Why it matters: Lovable 公开了每分钟十亿 token 规模下的多供应商负载均衡方案,可借鉴其用 PID 控制器和项目级粘性保住 prompt caching 的做法。

2/26Thu
2/23Mon
  1. OpenAI Developer Blog · Codex72

    用 Codex 跑 25 小时长时程任务:一份可复用的项目记忆文件栈

    OpenAI 用 GPT-5.3-Codex 在 Extra High 推理档下从空仓库连续运行约 25 小时、消耗约 13M token、生成约 3 万行代码,做出一个可测试的设计工具。

    Awaiting translation

    Why it matters: 作者用 25 小时、13M token 的实测展示长时程智能体如何靠持久化项目记忆和逐里程碑验证保持不跑偏。

2/5Thu
  1. Martin Fowler · Exploring Generative AI75

    Context Engineering for Coding Agents: A Look at Configuration Options, Using Claude Code as an Example

    A Martin Fowler team article breaks down context engineering for coding agents, sorting context configuration into reusable prompts (instructions and guidelines), context interfaces (tools, MCP Servers, Skills), and workspace files. It then splits these by "who decides what gets loaded" into three categories: the LLM, the human, and the agent software.

    Why it matters: Using Claude Code as an example, this piece walks through how to configure context for coding agents and lays out the trade-offs between loading on demand and building up gradually.

1/27Tue
  1. Martin Fowler · Exploring Generative AI70

    用 AI 智能体写代码时如何评估内部代码质量:CCMenu 加 GitLab 支持的实测

    Martin Fowler 用给 Mac 应用 CCMenu 增加 GitLab 支持的实验,考察 AI 智能体生成代码的内部质量。他先后用 Windsurf 加 Sonnet 3.5、Claude Code 加 Sonnet 4.5,让智能体参照现有 GitHub 的 API 封装、feed reader 和响应解析三个文件实现 GitLab 版本。

    Awaiting translation

    Why it matters: 作者用给 CCMenu 加 GitLab 支持的实测,展示 AI 智能体在内部代码质量上会引入哪些隐性技术债。