用「每次被采纳结果的成本」衡量编程智能体的开销
Original title: Measure coding-agent cost per accepted outcome
The title and summary in the selected language are awaiting translation.
作者认为按请求计费会掩盖重试、重复上下文和被拒补丁,真正该看的是总花费除以实际合并的改动,并把人工评审时间算进去。
Claude Code, GitHub Copilot, and similar agents bill for tokens, but the number that matters is cost per accepted outcome: total spend divided by the changes you actually merged, with human review time included. Per-request pricing hides retries, repeated context, and rejected patches.
An agent run can make several model calls, repeat a large context prefix, invoke tools, retry failures, and produce a patch that still needs review. Per-request token cost does not capture whether that run created value.
The useful unit is cost per accepted outcome: a merged patch, an accepted diagnosis, a reviewed migration plan, or another result your workflow defines before the run starts.
Measure cost per useful outcome
Tokens per request are necessary telemetry, but they are not an outcome. A request can be large and useful or small and discarded. Track provider cost beside acceptance and human effort.
At minimum, log fields like these. The values below are an example record, not a benchmark:
{
"workflow": "draft_pr",
"model": "provider-model-id",
"inputTokens": 84200,
"outputTokens": 6100,
"cachedInputTokens": 52000,
"costUsd": 2.41,
"status": "merged",
"humanReviewMinutes": 18
}That record lets you ask better questions. Which workflows are expensive but rarely accepted? Which cheaper models are good enough? Which prompts benefit from caching? Which tasks look cheap but create too much review drag?
Context volume is one controllable lever
Agentic systems often waste money by treating context as a landfill. More files, more logs, more docs, more examples. The model can handle it, so why not? Because every irrelevant token competes with relevant attention and may cost real money.
Context should be routed like data. Start small, retrieve what is needed, and expand only when the task justifies it.
Task
->
Classify workflow
Small context pack
enough?
Implement or answer
If evidence is thin, retrieve targeted files instead of dumping the whole repository.
For coding agents, a context pack might include the task, repo map, relevant feature note, nearby files, and test commands. It should not include the entire repository unless the workflow genuinely needs architectural search.
Prompt caching changes the math
Prompt caching is one of the most practical cost tools because agent workflows often reuse the same stable context: system instructions, repo conventions, API docs, tool descriptions, and feature notes. If your provider supports caching, arrange prompts so stable blocks remain stable. Put volatile task details later.
For a deeper walkthrough of cached token pricing, cache hits, cache misses, and prompt structure, read prompt caching for coding agents.
The shape looks like this:
[stable] agent role and review rules
[stable] repo map and commands
[stable] feature invariants
[semi-stable] relevant file excerpts
[volatile] user task
[volatile] latest tool outputsSmall formatting changes can break cache hits, so treat stable prompt sections like versioned assets. Do not casually regenerate them with different whitespace every run. Cache hit rates are not glamorous, but they compound fast at scale.
Route by job, not by brand
Do not assume one model is the best choice for every step. Treat routing as an experiment by job class:
| Job | Candidate to evaluate |
|---|---|
| Repo search and summarization | Fast, low-cost model |
| Boilerplate test generation | Fast, low-cost model |
| Risky architecture planning | Strong reasoning model |
| Final diff review | Strong reasoning model |
| Repetitive formatting | Deterministic tooling |
The trick is to avoid routing purely by vibes. Run evaluations on your own workflows. A cheaper model that handles 80 percent of routine test generation is a win, even if you still use a premium model for design review. A premium model that prevents one bad migration can also be a bargain.
Do not infer routing quality from price tier or product positioning. Evaluate candidate models on representative tasks, with the same tools, context, and acceptance rule.
Put a budget in the loop
Agents need budgets the same way services need timeouts. Without a budget, a stuck agent can keep expanding context, retrying, and asking a premium model to think harder about a task that should have been escalated to a human.
Budgets are also useful socially. They give teams a neutral way to stop a run without debating whether the agent is “close.” Once the run crosses the threshold, it has to produce evidence: what it tried, what it learned, and why more spend is justified.
An illustrative budget shape might look like this; choose thresholds from observed runs and your own risk tolerance:
type AgentBudget = {
maxCostUsd: number;
maxToolCalls: number;
maxContextTokens: number;
requireApprovalAboveUsd: number;
};
const draftPullRequestBudget: AgentBudget = {
maxCostUsd: 1.25,
maxToolCalls: 20,
maxContextTokens: 120_000,
requireApprovalAboveUsd: 0.75,
};The approval threshold is important. It lets a workflow continue when the task is worth it, but it forces the agent to explain why it needs more budget. That explanation alone can reveal a bad plan.
How do you calculate ROI with review time included?
The bill is not only provider cost. Human review and cleanup time are part of the workflow cost. Compare them in one unit instead of treating a cheap API call as automatically efficient.
Use a simple model first:
net_minutes_saved = estimated_manual_minutes
- human_review_minutes
- cleanup_minutes
- (agent_cost_usd / loaded_engineer_cost_per_minute)
cost_per_accepted_outcome = total_agent_cost_usd
/ accepted_outcomesThe estimate is imperfect, so keep the raw measurements as well. Track acceptance criteria, reviewer time, cleanup time, and why rejected runs failed.
Retries are the silent killer. One failed tool call can trigger another model call, which triggers more context, which triggers another attempt. Long-running agents need observability around loops: why did the agent retry, what changed, and when did it stop?
Generated tests can also multiply cost if every run asks the model to re-read the same failure output. Tooling should summarize failures and preserve stable context. Do not paste a 20,000-line test log into a premium model when the useful signal is a single assertion diff.
Another multiplier is unbounded chat history. Long conversations feel convenient, but old turns can become expensive and misleading. For repeatable workflows, compress history into structured state: task, decisions, files changed, checks run, open questions.
Put defaults in the workflow
A workable starting policy is a set of defaults embedded in the workflow:
- Start with the cheapest model that has passed evaluation for the job.
- Retrieve targeted context before expanding the window.
- Cache stable prompt blocks.
- Stop after budget thresholds and ask for approval.
- Log cost next to outcome, not in a separate dashboard nobody opens.
Developers will follow cost controls that make the right path easy. They will route around controls that feel like paperwork.
Spend where it buys judgment
The goal is not to minimize every agent run. Some tasks deserve the expensive model. Architecture review, security-sensitive changes, migration planning, and ambiguous incidents are places where better reasoning can be worth the premium. The waste is using that premium reasoning for tasks a cheaper model, a linter, or a deterministic script can do.
Token ROI gets easier when you stop treating the agent as one thing. It is a workflow made of steps. Each step has a job, a model choice, a context budget, a cache profile, and an outcome. Tune those pieces, and the bill becomes less mysterious.
Review the ledger on a schedule. Retire workflows that are rarely accepted, tighten those dominated by review time, and expand only where measured outcomes justify the additional cost.
Source: DevAgentStack · Field Notes · devagentstack.com