Addy Osmani 建议为 AI 智能体配备校验产出的手段,测试是其中一部分。他推荐四类测试:模拟真实用户流程的端到端测试作为基准真相;用基于属性的测试声明绝不应发生的情况并生成数千用例尝试突破约束;替换系统时用随机输入对比新旧版本;要求测试快速且确定性运行,否则智能体可能学会重试而非修复问题。
当前语言的正文正在等待翻译,暂时显示原文。
To make your agent produce high-quality output, equip it with a way to check its work. Tests are part of that.
Here are some of the ones I think are worth investing in.
- Write end-to-end tests that simulate real user flows and act as the ground truth
- Use property-based testing to specify what should never happen and generate thousands of test cases that attempt to break those constraints
- If replacing a system, compare the old and new versions using randomly selected inputs
- Require tests to run quickly and deterministically, which reinforces the agent to retry when the loop is slow or unreliable (the agent may learn to retry, not fix the problem)
Example tests say what should happen; properties say what must not.
来源:Addy Osmani · x.com