Skip to content
Original
Ryan Lopopolo·· 04/10/2026SelectedAI score65

怎样才算把活干好:写清非功能性需求才能让 AI 智能体收敛

Original title: What Does It Mean to Do a Good Job?

The title and summary in the selected language are awaiting translation.

AI overview

Ryan Lopopolo 认为,AI 让验证问题变得明显,因为每个真实任务都依赖一个我们几乎从不写下来的问题,即怎样才算把活干好。产出和评审都涉及语气、品味、风险容忍度、打磨程度、可接受的捷径和完成标准等大量非功能性决策,过去团队靠组织设计、社交规范、招聘和入职把这些隐含规则传递给人,而模型无法走招聘流程,因此交给它的任务基本都欠规范。

Why it matters

作者以在 OpenAI 做代码智能体的经历说明,非功能性需求不写下来,评审智能体就会陷入无休止的拉扯。

Full text

Ryan Lopopolo•April 10, 2026

Verification was always the problem. AI makes that obvious because every real task is downstream of a question we almost never write down. What does it mean to do a good job?

Producing an artifact and reviewing it both require hundreds of small decisions about non-functional requirements: tone, taste, risk tolerance, how much polish is enough, what shortcuts are acceptable, what counts as done. Historically, teams wrote almost none of that down. We encoded it in org design, social norms, hiring loops, onboarding, and repeated exposure to people who already knew the answer.

A cream-colored rectangular object sits clamped inside a precision inspection jig on a white background. The jig is rendered in brushed aluminum with teal-tipped stops, dial gauges, and small hanging calibration tags, all arranged around the object to suggest exacting measurement from multiple sides. The image is a clean, close isometric 3D render with soft studio lighting and a severe, minimal, object-first composition.

That worked well enough when the worker was a coworker. You could hire for judgment, build trust over time, and let the person absorb the unwritten rules. You do not get to run your AI through a hiring pipeline. And because models have seen trillions of documents that embody every possible permutation of those choices for nonfunctional requirements, basically every task we hand them is under-specified.

One concrete example from my work at OpenAI: for our code implementation agent and various reviewer agents, we had to write down choices human reviewers usually carry in their heads. Implementation agents must acknowledge review feedback, but they can accept and fix, accept and defer, or push back. Reviewer agents are told to bias toward merging and only surface P2s and above.

Without that guidance, the reviewers endlessly bully the implementer and nothing converges. Human reviewers usually know to unblock while still providing high-signal feedback. The models do not unless you say it.

This is why verification suddenly feels like the whole problem and why harness engineering is a productive area for applied AI engineering. The missing spec was always there. Humans were just better at smuggling it in through shared context.

Post-training optimizes these tools to be helpful assistants, not task-specific experts. The models crave text and they are rewarded for how well they follow instructions, so there is an inherent tension in RL between instruction-following fidelity and “creativity” in reasoning. If you want the models and agents to do a good job, write down what that means, then add nuance only when the coarse instruction starts to overfit.

Source: Ryan Lopopolo · hyperbo.la