如何设计可评审的编码智能体工作流
Оригинальный заголовок: Design a reviewable coding-agent workflow
Заголовок и краткое изложение на выбранном языке ожидают перевода.
作者提出可评审的编码智能体工作流应包含五个阶段:说明意图、提供仓库上下文、明确智能体可改动的边界、运行验证、要求以风险为重点的总结,跳过任一阶段只会把成本转移到评审环节。
Полный текст на выбранном языке ожидает перевода. Пока показан оригинал.
An AI coding workflow turns a request into reviewable work in five stages: state the intent, supply repository context, set explicit boundaries on what the agent may change, run verification, and require a risk-focused summary. Skipping any stage moves the cost to review time rather than removing it.
A successful demo proves that an agent completed one path in one environment. It does not prove that the same request is safe in a repository with hidden coupling, stale documentation, slow tests, and public contracts.
The missing ingredient is often not a longer prompt. It is a working environment where the agent can answer three questions before it acts: what matters, what is allowed, and what evidence counts as done?
From coder to orchestrator
Manual coding rewards local precision. You hold a lot in your head, make a change, run a tight feedback loop, and adjust. When you transition to agentic work, you are still accountable for the code, but your main leverage comes from shaping the work surface: context, constraints, tools, checkpoints, and review.
Consider a hypothetical request: “fix the auth redirect bug.” It does not say which redirect is wrong, whether session behavior may change, which files are in scope, or what reproduces the failure. An agent can satisfy one interpretation by rewriting the session module or changing a dependency while missing the intended one-line behavior.
Replace vague intent with an executable boundary:
goal: Reduce duplicate session validation logic in the checkout flow.
allowed_changes:
- src/checkout/**
- src/auth/session.ts
out_of_scope:
- database schema changes
- auth provider migration
acceptance_checks:
- npm test -- checkout
- npm run typecheck
review_focus:
- no behavior change for expired sessions
- preserve existing redirect URLsThis is not fancy. That is the point. The agent sees the work as a bounded operation instead of a mood. The reviewer sees the same thing. Future you has a breadcrumb trail for why the patch looks the way it does.
What does the basic agent loop look like?
An agentic workflow should make uncertainty visible early. Use a loop that requires inspection before editing and verification before summarizing.
Intent
->
Repo reconnaissance
Plan with boundaries
->
Small implementation
Run checks
->
Review diff and risks
If checks fail, diagnose the failure and loop back to the plan.
This loop is mundane, but it changes the failure mode. Without reconnaissance, the agent often optimizes for the first plausible file it finds. Without a bounded plan, it may refactor through unrelated layers. Without checks, it can produce code that reads well and fails quietly. Without a risk summary, it can leave the reviewer with the hardest job: discovering what changed by reading everything from scratch.
The orchestration skill is deciding how much friction to add. For a typo, the loop is one minute. For a payment flow, the loop should feel almost bureaucratic. That is not a lack of trust in the agent. It is respect for blast radius.
Give agents landmarks, not lore
Every team carries lore: why a helper exists, which tests are slow but important, which customer has a weird edge case, which module looks abandoned but is actually load-bearing. Agents do not inherit that lore unless we put it somewhere they can retrieve.
The first useful habit is creating landmarks. A landmark is a durable file or convention that answers a repeated question. It can be a short FEATURE.md, a package-level README, a test naming convention, or an ADR. The point is not documentation theater. The point is making the codebase navigable by a worker that starts every task with amnesia.
For example, a checkout feature note can be short:
# Checkout
Owner surface: src/checkout
Critical invariants:
- Never create an order before payment authorization succeeds.
- Preserve the original return URL through session refresh.
- Treat inventory holds as best-effort until payment capture.
Useful checks:
- npm test -- checkout
- npm test -- auth/sessionThis file does not guarantee a correct patch. It makes three invariants and two relevant checks discoverable before the edit, which gives both the agent and reviewer a concrete contract.
Be precise about freedom
The wrong response to agent mistakes is to remove all freedom. If you prescribe every line, you are back to typing through a slower keyboard. The better move is to be precise about where freedom is useful.
Separate tasks into three zones. The first is fixed: public contracts, data migrations, security invariants, compliance-sensitive behavior, and approved user-visible copy. The second is guided: internal structure, helper names, test organization, and implementation strategy. The third is exploratory: finding duplication, proposing approaches, and comparing tradeoffs.
That split gives the agent room to be valuable without pretending it already has all your project context. A direction like “you may propose a new helper, but do not change the API response shape” is far better than “refactor this.” It gives the model a lane.
Ask for evidence
Require the agent to attach evidence to the summary. “Implemented session validation cleanup” is not enough; the reviewer needs the behavior changed, checks run, relevant files, and remaining risk.
When done, summarize:
1. Behavioral change, if any.
2. Files changed and why.
3. Commands run with pass/fail results.
4. Known risks or follow-up checks.This is where agentic workflow optimization becomes less about speed and more about compression. A good agent run compresses exploration, implementation, and verification into a reviewable unit. A bad run produces a pile of plausible code and a cheerful paragraph.
Does a larger prompt fix bad results?
Adding more prose does not resolve an undefined product decision, a hidden invariant, or an unknown verification command. It can make the request longer while leaving the same ambiguity intact.
Use a better contract instead: smaller scope, explicit invariants, discoverable repository landmarks, and named checks. Ask the agent to inspect first and explain residual risk last.
That is the real art of giving directions. You are not trying to dominate the model. You are trying to design the room it works in, so the obvious next action is also the responsible one.
The developer does not disappear in that setup. A human still owns product decisions, risk acceptance, and the merge. The agent can compress exploration, editing, and verification only when the surrounding workflow makes those responsibilities visible.
Источник: DevAgentStack · Field Notes · devagentstack.com