不稳定测试:重跑变绿对智能体意味着什么
Original title: Flaky tests: what a green rerun tells an agent
The title and summary in the selected language are awaiting translation.
文章用概率模型算出,失败率 1% 的测试要跑 299 次才有 95% 概率被发现,而单次重跑有 99% 概率变绿,因此重跑通过几乎不提供信息。作者建议智能体循环中不要用重跑覆盖原始失败记录,并把已知不稳定测试单独标记,例如用 pytest-rerunfailures 的 @pytest.mark.flaky(reruns=n) 按测试标注,而不是全局 --reruns 3。
A test failing 1% of the time needs 299 runs to reveal itself, and its rerun is green 99% of the time. The arithmetic, a dataset, and what it means for agent loops.
An agent finishes a change, runs the suite, and one test is red. It reruns that test. The test passes. The agent reports the work is done.
That sequence is reasonable, common, and almost entirely uninformative, and you can see why with one line of arithmetic. This page works out how much a rerun is worth, how many runs it takes to notice a flaky test at all, and what both numbers imply for how you let an agent decide it is finished. Every figure below comes from a script that ships with the article, and each one is computed two independent ways.
The arithmetic
Model a flaky test as a coin: on each independent run it passes with probability p. It is flaky whenever p is strictly between 0 and 1. Now run it n times. You only learn the test is flaky if you see disagreement — at least one pass and at least one failure. If every run agrees, the test looks deterministic and nothing draws your attention to it.
P(all n runs pass) = p**n
P(all n runs fail) = (1-p)**n
P(you see a flake) = 1 - p**n - (1-p)**n
That is the whole model. The shipped script evaluates it directly and also simulates it with 200,000 Monte-Carlo trials per cell from a fixed seed; the largest gap between the closed form and the simulation across all 48 cells is 0.0028, which is sampling noise at that trial count. The table below is the closed form.
| Failure rate | 2 runs | 3 runs | 5 runs | 10 runs | 20 runs | 50 runs | 100 runs |
|---|---|---|---|---|---|---|---|
| 1% | 2.0% | 3.0% | 4.9% | 9.6% | 18.2% | 39.5% | 63.4% |
| 5% | 9.5% | 14.3% | 22.6% | 40.1% | 64.2% | 92.3% | 99.4% |
| 10% | 18.0% | 27.0% | 40.9% | 65.1% | 87.8% | 99.5% | 100.0% |
| 20% | 32.0% | 48.0% | 67.2% | 89.3% | 98.8% | 100.0% | 100.0% |
| 30% | 42.0% | 63.0% | 83.0% | 97.2% | 99.9% | 100.0% | 100.0% |
| 50% | 50.0% | 75.0% | 93.8% | 99.8% | 100.0% | 100.0% | 100.0% |
Read the top row. A test that fails one run in a hundred — the sort of thing that exists in every suite of any size — reveals itself in ten runs about 9.6% of the time. Run it a hundred times and you still miss it more than a third of the time.
Turn that around and ask how many runs you need before you would probably have noticed:
| Failure rate | Runs for a 95% chance of ever seeing it |
|---|---|
| 1% | 299 |
| 5% | 59 |
| 10% | 29 |
| 20% | 14 |
| 30% | 9 |
| 50% | 6 |
Two hundred and ninety-nine runs for the 1% test. If your CI runs the suite twenty times a day, that is a fortnight of executions before you are more likely than not to have seen the thing twice behave differently — and only if someone is looking across runs rather than at each run alone.
What a green rerun is worth
Now the case that matters for agents, because it is the one they hit constantly. A test has just failed. You rerun it once. What does a green rerun tell you?
The rerun passes with probability p. For a mildly flaky test p is close to 1:
| Failure rate | Chance the single rerun is green |
|---|---|
| 1% | 99.0% |
| 5% | 95.0% |
| 10% | 90.0% |
| 20% | 80.0% |
| 30% | 70.0% |
| 50% | 50.0% |
At a 1% failure rate the rerun is green 99 times out of 100. That is the number to sit with, because it is the same 99% whether the original failure was pure noise or the first sighting of a real intermittent bug that only manifests one run in a hundred. The rerun does not distinguish between those two cases. It cannot: both hypotheses predict a green rerun with overwhelming probability.
This is why "it passed on retry" is such a weak piece of evidence, and why an agent that treats a passing rerun as proof of correctness has, in information terms, learned close to nothing. The failure told you something. The rerun told you almost nothing, and then it deleted the thing the failure told you.
The asymmetry that makes it worse for agents
A human running into a red test carries context an agent usually does not: whether this test has been annoying for months, whether the failure message looks like a timeout or like an assertion on a value, whether anybody else has seen it. That context is doing the real work when a person says "that one's flaky, ignore it."
An agent in a fresh session has none of it. What it has is a red result, a rerun, and a green result — the exact sequence the table above says is nearly uninformative. And agents are structurally biased toward the reassuring reading, because a green suite is a terminating condition and a red one is more work.
Two habits follow from the arithmetic, and neither is complicated:
Never let a rerun overwrite the record of the failure. The failure is the evidence; the rerun is not a correction to it. If your agent's loop reruns a failing test, the transcript should still carry the original failure, its message, and the fact that a retry was used. A session that ends "all tests pass" after a silent retry has destroyed the only signal it had.
Make the flaky ones a separate category, not a judgement call. A test that is known-flaky should be marked as such in the code, not diagnosed fresh each session by whatever is running the suite. pytest-rerunfailures exists precisely for this: the pytest-dev plugin, which describes itself as re-running "tests to eliminate intermittent failures", lets you mark individual tests with @pytest.mark.flaky(reruns=n) rather than applying a blanket retry to the whole suite. Its README (read 2026-09-21; the repository stood at 477 stars, with its most recent commit 6f843b6098c6 dated 2026-09-18) states the requirements plainly: Python 3.10+ or PyPy3, and pytest 8.2 or newer.
The distinction matters more than the plugin. A per-test marker is a claim someone made deliberately and can be audited; a global --reruns 3 is a policy that quietly applies the 99%-green-rerun effect to every test you own, including the ones that were about to tell you something.
Using the numbers deliberately
The same table that makes reruns look bad makes repetition look good, if you do it on purpose and at the right moment. Detecting a 10% flake takes 29 runs for a 95% chance; that is seconds for a fast unit test. So:
- When a test fails once and you suspect flakiness, do not rerun it once. Run it the number of times the table says you need for the failure rate you suspect, and record how many passed. Twenty runs of a supposedly-1%-flaky test that come back twenty green is weak evidence; twenty runs that come back three red is a measured failure rate, and now you have a number instead of an argument.
- When you add a test that touches time, ordering, the network, or shared state, run it fifty times before you trust it once. That is cheap at the point of writing and expensive to retrofit after it has been failing intermittently in CI for a month.
- Track the failure rate rather than the failure. A test's flake rate is a number you can watch move. A list of tests that "sometimes fail" is not.
What this does not cover
The independence assumption is the weak point. Real flakes are frequently not independent coin flips. They cluster: a test fails when the machine is loaded, or when it runs after a particular other test, or at a particular time of day. Order-dependent failures in particular break the model completely — a test that fails only when it runs second will fail 100% of the time in one ordering and 0% in another, and no number of reruns in a fixed order will ever reveal it. Treat the tables as the best case for detection by repetition.
The failure rates here are illustrative, not measured. The 1%, 5% and 10% columns are round numbers chosen to span the range people talk about; they are not an empirical distribution of real-world flake rates, and this page does not claim to know what that distribution is. What is computed, and what ships in the dataset, is the detection probability given a rate.
No agent was measured for this. The claims about how agents behave in a rerun loop are arguments from how the loops are constructed, not from an experiment on transcripts. The arithmetic is the part that is verified.
One plugin, named because it is the one I checked. pytest-rerunfailures is cited here with the figures I read on 2026-09-21. There are other reruns plugins and other test runners with equivalent features; this is not a survey and no comparison is implied.
Run it yourself
The script is standard library only and takes about a minute at the default 200,000 trials per cell:
python flaky-test-rerun-math.py
python flaky-test-rerun-math.py --trials 1000000 # tighter simulation
It prints all three tables, reports the largest gap between the closed form and the simulation, and writes datasets/flaky-test-rerun-math.csv with 48 rows: the pass rate, the failure rate, the run count, the closed-form detection probability, the simulated one, and the absolute difference between them. If you change RATES or RUNS at the top, the tables and the dataset follow.
The conclusion I would defend is narrow. A green rerun is not evidence that the code is correct; it is evidence that you ran a probably-flaky test a second time. Whether that distinction matters in your suite depends on how many nearly-deterministic tests you own — and the first row of that first table is a good argument for finding out.
Code and data
- flaky-test-rerun-math.py — the complete listing used in this article.
- flaky-test-rerun-math.csv — the data behind the numbers here.
Sources
Source: Vibe Code Textbook · Articles · vibecodetextbook.com