编码智能体的表现取决于 harness,而非只取决于模型
Original title: Your coding agent is only as good as its harness.
The title and summary in the selected language are awaiting translation.
作者认为编码智能体的 harness 是模型外的软件层,负责组装上下文、暴露工具、控制运行循环、执行权限并评估结果,因此同一模型在不同 harness 下表现会完全不同。
A coding agent harness is the software layer around the model: it assembles context, exposes tools, controls the run loop, enforces permissions, and evaluates results. The model writes the code, but the harness decides what the model sees, what it can touch, and when it stops.
Every few months, developers ask the same question: which model is best for coding? It is a useful question, but it is no longer the whole question. A strong model inside a weak workflow will still miss files, call the wrong tools, loop too long, stop too early, or produce a diff nobody trusts. That workflow layer is the harness.
Consider a failure mode: a team switches models expecting cleaner pull requests, but the workflow still has no scope boundary, targeted retrieval, or verification requirement. A more capable model can still produce a larger out-of-scope diff. Changing the model did not repair the workflow around it.
The VS Code team’s post on the coding harness behind GitHub Copilot names this layer well. The harness decides what the model can see, what it can do, how long it can keep working, and how the result is judged.
If you work with coding agents, you already have a harness. It might be an editor integration like Cursor or Devin Desktop, a terminal agent like Claude Code or OpenAI Codex CLI, or Copilot in VS Code. It might be a custom script wrapped around a model API or an issue-to-PR bot wired into CI. Designed or not, the harness is there.
What is a coding agent harness?
A model produces text. A coding agent does work.
The harness is the layer that turns one into the other. It gathers context, exposes tools, applies edits, runs commands, tracks state, decides when to continue, and packages the result for review.
The agent loop: context, execution, and self-correction
Most modern coding agents follow a continuous autonomous loop managed by the harness:
- Assemble Context: The harness collects the repo map, instruction files (
AGENTS.md), open buffers, and tool definitions. - Model Reasoning: The model evaluates the task and selects an action (reading a file, searching symbols, editing a block, or running a shell command).
- Tool Execution & Guardrails: The harness intercepts the call, checks permission boundaries, runs the tool in a sandbox, and captures stdout, stderr, or diffs.
- Observation & Self-Correction: When a command or test fails, the harness feeds the exact error back to the model, allowing it to self-correct its implementation rather than aborting.
- Stopping Conditions: The loop terminates when verification succeeds, the task goals are met, or safety boundaries (max turns / timeout) trigger.
In practical terms, a harness answers questions like these:
- What instructions does the agent see before the task starts?
- Which files, docs, memories, and prior messages are included?
- Which tools can the model call?
- Can those tools read, write, execute, browse, or only propose changes?
- How are file edits applied?
- What happens when a command fails?
- When does the loop stop?
- What evidence proves the answer is useful?
That is a lot more than prompting.
A prompt says what you want. A harness shapes the entire operating environment around the request.
Task
->
Context
Model
Tools
loop
Verification
The harness decides what enters the loop and what counts as done.
This is why two products can use the same model and feel completely different. The model may be similar. The harness is not.
Why do harnesses matter now?
Model benchmarks and context windows describe only part of an agent product. The VS Code and Claude Code sources linked below also document architecture around tools, task state, permissions, context, and the loop around the model. Reading both makes the comparison concrete without pretending the products expose identical internals.
That shift is overdue.
Claude Code documents a terminal-oriented agent loop. Copilot in VS Code is editor-native and integrates with editor context, tools, and review surfaces. Cursor makes its own editor-level choices. Cloud and CI agents make a different choice again: queued work runs away from the interactive editor and returns artifacts for review.
The useful comparison is not which one is “the” harness. The useful comparison is what each harness optimizes for.
- Terminal-first harnesses optimize for directness, scriptability, and command-line workflows.
- Editor-native harnesses optimize for IDE context, language services, visible diffs, and interactive review.
- Cloud or CI harnesses optimize for repeatability, queued work, and integration with pull request workflows.
- Custom internal harnesses optimize for company-specific tools, permissions, and evaluation.
The architectural choices matter because they decide how much trust the agent can earn.
Which harness choices do users actually feel?
Most developers do not need to know the term “agent loop.” They feel it when the agent makes a review mess, or doesn’t. Harness design shows up in that feeling.
Context assembly
Context is the first product decision. What does the model see, and in what order?
A good harness separates stable context from volatile context. Stable context includes system rules, repo maps, tool schemas, coding conventions, and review expectations. Volatile context includes the current task, retrieved files, tool output, diffs, and errors.
That separation improves quality and cost. Stable context can be written once, reviewed, reused, and sometimes cached. Volatile context can stay focused on the current job. This is the same discipline behind prompt caching for coding agents: stable context is valuable when it stays stable.
A weak harness turns context into soup. It mixes old chat turns, random file excerpts, changing summaries, huge logs, and task instructions into one long prompt. The model may still produce something plausible, but the workflow becomes harder to debug. When the answer is wrong, nobody knows whether the failure came from the model, the retrieval layer, stale instructions, or irrelevant context crowding out the important part.
Tool exposure
Tools are not just capabilities. They are permissions and product surfaces.
An agent with no tools has to guess from chat context. An agent with too many tools can wander. A useful harness exposes the smallest set of tools that helps the model gather evidence, change the right files, and verify the result.
That is why tool names and schemas matter. search_docs is different from call_internal_api. get_schema(tableName) is different from run_sql(sql). create_draft_release_note is different from post_to_status_page.
This is the same design pressure behind giving agents a pair of hands and not giving agents database access too early. A harness should reduce guessing before it expands authority.
Good tool exposure has three properties:
- The model can tell when the tool is relevant.
- The tool has a narrow, typed input shape.
- The result is structured enough to support the next decision.
The worst version is a tool buffet: arbitrary shell, arbitrary SQL, broad write access, vague tool descriptions, and no review boundary. That may look powerful in a demo. It is exhausting in production.
Loop control
The loop is where agents become useful and expensive.
A coding agent rarely answers once. It reads, plans, edits, runs a command, sees a failure, edits again, runs a narrower check, and summarizes the result. That loop needs limits.
Good harnesses have budgets. At a minimum:
- A cap on tool calls per run.
- A cap on cost or token spend per run.
- A cap on runtime before the agent has to check in.
- An approval step for risky tools, not just risky outputs.
Budgets are not just cost controls. They are trust controls. Without them, a stuck agent can keep spending tokens and touching more files while the human waits. With them, the agent has to surface what it tried, what it learned, and why it needs more room.
Loop control also includes stop behavior. An agent should not stop because it has a nice explanation. It should stop because it has met the task’s evidence bar or because it cannot proceed honestly.
That distinction matters. “I changed the file” is not done. “I changed the file, ran the relevant build, and here is the remaining risk” is closer.
Edit application
File editing is a harness feature, not just a model skill.
Some harnesses ask the model to output patches. Some use structured edit tools. Some apply search-and-replace operations. Some let the model write whole files. Each choice changes the failure mode.
Patch-based editing can be reviewable and compact, but it may fail when context drifts. Whole-file editing can be simple, but it risks accidental churn. Search-and-replace can be precise, but brittle when the match is ambiguous. IDE-native edits can use language context, but depend on the environment and extensions available.
The practical question is not which edit style sounds smartest. It is whether failed edits are observable and recoverable.
When a harness applies an edit, it should be clear:
- Which file changed?
- What exact diff was produced?
- Did the edit apply cleanly?
- Did formatting or generated output change unrelated files?
- Did the agent inspect the result before claiming success?
This is where harness quality becomes review quality. A good harness makes it easy for the human to see what happened.
Evaluation
The VS Code post is especially strong on evaluation because it treats harness changes like product changes. If a system prompt, tool definition, or loop behavior changes, the team wants evidence that agent behavior improved rather than merely changed.
A team does not need a large benchmark suite to begin measuring its workflow. It does need repeatable tasks and an acceptance rule.
Start with a small eval set:
- One bug fix in a familiar module.
- One test generation task.
- One refactor across two or three files.
- One docs update tied to code behavior.
- One task the agent should refuse or narrow before acting.
Run those tasks when you change repo instructions, tool access, model defaults, or workflow prompts. Track whether the agent reached the right outcome, how many tool calls it used, how much review cleanup it created, and whether the final summary was honest.
This does not have to be fancy. A spreadsheet is better than vibes. A repeatable script is better than a spreadsheet. A real eval harness is better still.
The important move is measuring the workflow, not worshiping a benchmark number.
What can teams control today?
Teams using an existing agent product still control important parts of the local operating environment.
Your repo and workflow already supply much of the harness. The agent reads your docs, calls your scripts, follows your instructions, discovers your tests, and produces diffs for your review process. If those pieces are messy, the agent has to compensate. Sometimes it can. Often it guesses.
The parts a normal team can improve are concrete:
- Root-level instructions:
AGENTS.md,copilot-instructions.md, or the equivalent for your agent environment. - Repo map: where important behavior lives and which entry points are canonical.
- Verification commands: the fastest useful build, test, lint, and content checks.
- Risk boundaries: auth, billing, migrations, production data, public APIs, SEO, analytics, deploy scripts.
- Tool policy: which tools are read-only, which propose changes, and which require approval.
- Prompt structure: stable context first, volatile task state later.
- Review contract: what the final answer must include before a human inspects the diff.
- Eval tasks: small repeatable tasks that reveal regressions in agent behavior.
This is why making a repo agent-ready is a harness problem. The repo is part of the runtime. The scripts are part of the runtime. The docs are part of the runtime. The review checklist is part of the runtime.
An agent-ready repo is not only easier for a model to understand. It is easier for any harness to operate safely.
A rough maturity picture
Use a short ladder to describe the workflow you actually operate. At the bottom, the agent has repository access but no defined verification step. In the middle, it has scoped tools, documented commands, and draft-only writes. A more mature workflow records acceptance, rejection, cost, and cleanup evidence and reviews changes to prompts or tools. A productized internal harness is another category, not a required destination.
The useful move is not climbing the ladder for its own sake. It is knowing which rung you are actually on before you hand the agent bigger work. The scorecard below is how I check.
The scorecard I would use
Here is the quick audit I would run before trusting an agent workflow with real engineering work.
Context
- Can the agent find the repo map without being told?
- Are stable instructions short, current, and enforceable?
- Are risky surfaces named?
- Are task-specific details kept separate from global rules?
Tools
- Does every tool have a clear job?
- Are read tools preferred before write tools?
- Are broad tools gated by approval?
- Are tool results structured and redacted where needed?
Loop
- Is there a budget for tool calls, cost, or runtime?
- Does the agent know when to stop and ask?
- Are failures summarized rather than hidden behind another retry?
- Can the human cancel or narrow the run cleanly?
Verification
- Are relevant commands documented?
- Does the agent run the smallest useful check?
- Does the final summary distinguish passed, failed, and not run?
- Are generated diffs easy to inspect?
Evaluation
- Do you have a small repeatable eval set?
- Do you track accepted work, rejected work, and cleanup time?
- Do prompt or tool changes get tested against realistic tasks?
- Can you tell whether a new model improved the workflow or only changed the style?
If the answer is mostly no, changing models may still help. It just will not fix the system around the model.
The trap of treating the harness as invisible
The easiest mistake is to treat harness behavior as something the vendor owns completely. Some of it is. Copilot in VS Code, Claude Code, Cursor, Codex, and other tools make deep product decisions you cannot directly change: tool protocols, edit strategies, prompt assembly, memory behavior, permission flows, model defaults, and UI surfaces.
But the vendor harness meets your local harness.
Your repo instructions, scripts, tests, docs, package structure, branch policy, CI gates, MCP servers, and review norms all become part of the agent’s operating environment. The VS Code post frames the model as the engine and the harness as the car. The extension of that picture is simple: your repo is the road. Without lanes, signs, and an inspection station, even a good car drifts.
That is good news. It means teams are not helpless while the model market shifts.
You can make the next agent run better by doing very ordinary engineering work:
- Delete confusing duplicate entry points.
- Add the missing test command.
- Write the invariant beside the feature.
- Tighten tool permissions.
- Log cost next to outcome.
- Keep stable context stable.
- Run the same five eval tasks before changing defaults.
None of that requires waiting for a new frontier model.
The senior engineer version
When someone asks which coding model is best, the honest answer is: best inside what harness, for what task, under what budget, with what tools, and judged by what evidence?
That may sound less satisfying than a leaderboard. It is also closer to how real software work happens.
The model is still important. Better reasoning, longer context, lower latency, stronger tool use, and cheaper cached tokens all matter. But model capability only becomes engineering value when the harness can route it toward the right work and prove the result.
For teams, the practical takeaway is simple:
Design the harness you already have.
Make context intentional. Make tools narrow. Make loops bounded. Make verification visible. Make evaluation repeatable. Then model upgrades become easier to judge because the workflow around them is stable enough to measure.
The future of coding agents will not be decided by models alone. It will be decided by the systems that let those models do useful work without leaving humans to clean up the mystery afterward.
For the repo side of harness design, start with making a repo agent-ready and the templates page. The development-environment guide maps the source, runtime, network, credential, and teardown boundary around that harness.
Source: DevAgentStack · Field Notes · devagentstack.com