Skip to content

Practice and tutorials

Best practices, reproducible workflows, and tutorials with commands and configuration examples.

Latest curated items

Items 41–60 · 117 total
8/19Wed
  1. Cline · Blog71

    Cline's Evaluation Methodology and Trace for Open-Weight Models

    Cline has open-sourced its evaluation methodology for open-weight models, along with a hill-climbing score and trace worth over one thousand dollars, available for download and analysis. The post lays out five heuristics from the Hill Climber's Checklist: set a North Star metric, quantify noise, break down failure modes by task/model/vendor, don't assume more thinking is always better, and keep a private evaluation set.

    Why it matters: Cline shares its evaluation methodology and a trace worth over a thousand dollars, offering five transferable hill-climbing heuristics that teams building their own harness can reference.

  2. OpenAI Developer Blog · Codex71

    OpenAI open-sources the Codex harness and the app-server client protocol

    OpenAI has open-sourced the harness that drives the Codex app, CLI, and IDE extensions, and through the Codex app-server client protocol it exposes capabilities like creating threads, starting turns, receiving events, and handling approval requests.

    Why it matters: With the Codex harness and app-server protocol now public, developers can see how to embed the agent in their own products and where the boundaries are.

8/18Tue
  1. Lovable · Blog74

    Lovable 如何把 lovable.dev 从 Next.js 迁到自家 TanStack Start 栈

    Lovable 用六个月把月访问 4200 万的 lovable.dev 从 Next.js 迁到自家 TanStack Start 托管栈,迁移期间两套框架并行运行,由代理 worker 按路由和用户分流,最终 Next.js 专属代码只占 3%。

    Awaiting translation

    Why it matters: Lovable 官方复盘把 90 万行代码从 Next.js 迁到自家 TanStack Start 栈,双框架并行、AI 智能体批量迁移和一次 OOM 事故的细节都可直接借鉴。

8/14Fri
  1. Addy Osmani · Blog78

    Addy Osmani’s loop engineering in practice: managing parallel agents with Claude Code’s goal and loop primitives

    Addy Osmani describes how he runs 5 to 10 agents in parallel every day, usually with at most 5 running at once, and breaks loop engineering into two Claude Code primitives: goal, which drives a single bounded task toward a verifiable definition of done, and loop, which reruns a prompt at fixed intervals, much like cron.

    Why it matters: The author breaks his daily routine of running 5 to 10 agents in parallel into two primitives—goal and loop—and shares reusable validation Skills and ways to write stop conditions.

8/11Tue
  1. Paper Compute · Engineering Blog80

    如何降低 AI Agent 成本:审计 500 个 Claude Code 会话后的三类成本泄漏

    作者用 tapes 记录并导出团队最近 500 个 Claude Code 会话,按工具名和参数做指纹统计,发现同一会话内完全重复的工具调用只占 3.6%(932/25587),真正的开销在别处。

    Awaiting translation

    Why it matters: 作者审计 500 个 Claude Code 会话,把成本拆成会话内、会话间与会话周边三类,给出可复用的排查方法。

8/7Fri
  1. InfoQ · AI Coding Presentations83

    How Spotify uses the backend coding agent Honk to continuously rewrite its codebase

    Spotify's platform team shared how the backend coding agent Honk evolved: it started out replacing migration scripts, then gradually took on build and test validation, raising merged PRs from 1000 per 3 months to 1000 per 10 days.

    Why it matters: Spotify walks through Honk's evolution from script-based migration to a backend coding agent, focusing on how validation and standardization determine whether generated code can be merged.

8/3Mon
  1. OpenAI · Codex Cookbook71

    Iterating on a Development Workflow with Codex: From AGENTS.md to Phased Build Files

    The OpenAI Codex Cookbook lays out a set of repository conventions for wiring Codex into your development process: use AGENTS.md for persistent repository instructions, PLANS.md as the source of phase plans, and split the work into phased build files under harness/build/, with each phase spelling out its goals, acceptance criteria, boundaries, and approval gates.

    Why it matters: OpenAI lays out a complete directory convention for constraining Codex with AGENTS.md, PLANS.md, and phased build files—one you can adapt to your own repository.

7/30Thu
  1. Martin Fowler · Exploring Generative AI78

    重构的经济收益:一次用 Claude Code 量化 token 节省的实验

    Martin Fowler 用一个约 15 万行、几乎全由 Claude Code 和 Cursor 写成的 Rust 应用做实验,把 17,155 行的数据访问层按严格重构步骤拆分,每步后用全新子智能体执行同一个代表性改动并记录 token 消耗。

    Awaiting translation

    Why it matters: 作者用同一改动反复跑重构前后对比,量化出重构对智能体 token 消耗的实际影响,并指出节省来自文件切分而非代码总量下降。

7/26Sun
  1. Hacker News · Context Engineering 讨论82

    Anthropic publishes new context engineering rules for Claude 5

    Thariq Shihipar, a member of Anthropic's engineering team, wrote up the new context engineering rules for Claude 5, saying the team has cut over 80% of the system prompt from Claude Code for models like Claude Opus 5 and Claude Fable 5, with no measurable loss on coding evals.

    Why it matters: Anthropic lays out the new context engineering rules for Claude 5 and explains how to trim the system prompt, CLAUDE.md, and Skills.

7/25Sat
  1. Cline · Blog79

    Cline 用递归自我改进让智能体自主把 Kimi K3 在 Terminal-Bench 2.1 提到 88.8%

    Cline 用一条提示词启动 17 小时连续运行的编码智能体,以 GPT-5.6-Sol 为 leader 模型,把 Kimi K3 在 Terminal-Bench 2.1 上的成绩从基线 69/89(77.5%,$79)提升到确认运行 79/89(88.8%,$49.8),超过 Moonshot 自报的 88.3%。

    Awaiting translation

    Why it matters: Cline 用一条提示词让智能体自主完成 17 小时爬山,把 Terminal-Bench 2.1 从 77.5% 提到 88.8%,可看具体修了哪些 harness 问题。

7/24Fri
  1. Lovable · Blog71

    Lovable 如何用 AI 黑客智能体集群对自己打夺旗赛

    Lovable 在内部搭建了一套进攻性安全程序,让 AI 智能体集群像人类攻击者一样探测系统入口,直到拿到可验证的漏洞证据。它用夺旗赛的思路做验证:把 flag 散布在基础设施和权限最高的产品界面中,不预埋任何漏洞,智能体取到 flag 就说明找到了真实入侵路径,而不是模型猜测。

    Awaiting translation

    Why it matters: Lovable 公开了用夺旗机制验证漏洞的内部攻防智能体编排方法,可迁移到自家安全测试流程。

7/22Wed
  1. Augment Code · Blog62

    What is loop engineering, and how are leading software engineering teams using it?

    Augment Code proposes loop engineering: designing agent loops that run from trigger to execution to validation to outcome, with agents handling the intermediate steps and humans stepping in only at checkpoints that require judgment. The article compares loop engineering with prompt engineering and context engineering as distinct layers, lays out five stages—trigger, execution, validation, outcome, and improvement—and describes four team-level loops already running in production: code review, ticket-to-PR, vulnerability remediation, and incident response.

    Why it matters: Augment Code breaks loop engineering into five stages—trigger, execution, validation, outcome, and improvement—and lays out four team-level loop patterns already running in production.

7/20Mon
  1. Addy Osmani · Blog74

    Addy Osmani on software factories: the visible factory and the hidden factory, where validation is the bottleneck

    Addy Osmani proposes that a software factory has three layers—loop, harness, and factory. The factory isn’t a smarter agent; it’s multiple loops with harnesses feeding into a single review gate, with humans controlling the outer loop.

    Why it matters: The author breaks the software factory into three layers—loop, harness, and factory—and points out that validation, not generation, is the real bottleneck.

7/19Sun
  1. Hacker News · Claude Code 高分80

    把闲置 Mac 配成 Claude Code 可完全控制的常驻机器:分步指南

    作者 ykev 发布一份分步指南,教用户把闲置 Mac 变成 Claude Code 可完全控制、开启 computer use 的常驻机器,可从手机 Claude app 或主 Mac 经 SSH 操作。

    Awaiting translation

    Why it matters: 作者把闲置 Mac 改造成 Claude Code 常驻控制机,给出从 SSH、免密 sudo 到 computer use 的完整步骤,可迁移到任意两台机器。

7/17Fri
  1. Ryan Lopopolo71

    Code Red needs a maintenance loop: use Codex /goal to turn emergency fixes into ongoing operations

    Drawing on his experience during Stripe’s first code yellow, the author points out that after most code reds, all that’s left is a post-mortem and exhausted engineers, while the metrics go back to being unowned. He argues that a code red should leave behind a maintenance loop, and that OpenAI Codex’s /goal command can turn a one-off coding request into an ongoing objective with clear completion criteria, letting a persistent cluster of agents continuously watch metrics, generate interventions, and request human review.

    Why it matters: Based on his Stripe code yellow experience, the author proposes using Codex’s /goal to turn one-off emergency fixes into a long-term maintenance loop that can carry over to SLO governance.

7/8Wed
  1. AI Hero · Skills Updates65

    AI Hero skills repo ships v1.1: adds /wayfinder, renames /to-spec and /to-tickets

    AI Hero's skills repo ships v1.1, renaming /to-prd to /to-spec, merging /to-plan and /to-issues into /to-tickets, and adding new Skills like /wayfinder, /research, and /prototype.

    Why it matters: The author walks through the full Skill flow from grilling to deployment and gives the migration commands for the renames, the merge, and the new /wayfinder—useful for anyone building an AI development workflow.

  2. Martin Fowler · Exploring Generative AI71

    在本地小模型上做智能体编码的实测体验

    Martin Fowler 在 M3 Max 48GB 和 M5 Pro 64GB 上实测 Qwen3.6 35B MoE、Gemma 4 31B/26B、Qwen Coder Next 80B 等本地小模型的智能体编码能力,按内存、速度、工具调用、代码正确性、上下文、任务复杂度、代码质量逐层筛选。

    Awaiting translation

    Why it matters: 作者用两个具体任务对比多款本地小模型的智能体编码表现,并给出可复用的任务选择标准。

7/6Mon
  1. 宝玉80

    The Claude Code team explains the four types of AI agent loops and how to use them

    The Claude Code team defines an AI agent loop as repeatedly running a work cycle until a predefined stopping condition is met, and splits it into four types—turn-based, goal-oriented, time-based, and proactive—along four dimensions: what triggers it, how it stops, the underlying instructions it uses, and the tasks it fits.

    Why it matters: The Claude Code team breaks loops into four types—turn-based, goal-oriented, time-based, and proactive—and lays out how each is triggered, how it stops, and how tokens are controlled.

7/5Sun
  1. Jesse Vincent68

    Developing Sen 2.0 by having Claude Code and an agent on Slack review each other's work

    At Prime Radiant, author Jesse Vincent used Claude Code—working through the Slackline command-line Slack client—to collaborate with his own agent Ada: Claude proposes changes, Ada reviews and tests them, then Claude deploys the updates, forming a development loop where the agents review each other.

    Why it matters: By looping two agents through mutual review, testing, and deployment, the author shows a transferable model for collaborative agent-based development.

7/3Fri
  1. Lovable · Blog88

    花掉 8.5 万美元 token 后,我在 Lovable 扩展智能体编程的经验

    Lovable 一名工程师从今年 1 月到 6 月把个人 token 花费从每月约 600 美元推到 5 月的约 2.5 万美元、累计约 8.5 万美元,同时把每周合并 PR 数从 20-30 个提升到 150 个以上。

    Awaiting translation

    Why it matters: 作者公开了自己每月约 2.5 万美元 token 的智能体开发配置,包括风险分级、多智能体评审和上下文管理,可迁移到其他团队。