Why Flaky Tests Win: Fixing the Fifth-Run Failure Pattern
A test that passes on the fifth run is not a nuisance; it is a signal. Flaky tests hide real risk, inflate CI cost, and train teams to ignore red builds until the one that matters slips through.
Nesqual Tech AI
A test that passes on the fifth run is not “mostly fine.” It is a liar with a green badge. In one enterprise CI system, a single flaky suite added 18 minutes to every merge request, burned 14 engineer-hours per week on reruns, and still let a production regression escape because everyone assumed the failure was random.
That is why flaky tests win: they teach teams to distrust the signal, not the software. Once a pipeline becomes noisy, engineers stop reacting quickly, release managers widen tolerances, and real defects hide behind the noise.
Why flaky tests outperform your best intentions
A flaky test wins by exploiting human behavior, not technical complexity. The first time it fails, you investigate. The third time, you rerun. The fifth time, you create a mental model: "this test is unreliable." That label is dangerous because it spreads.
The hidden economics of reruns
Consider a platform team running 2,400 test jobs per day across microservices, mobile, and data pipelines. If 2.5% of those jobs are flaky and each rerun costs 90 seconds of compute plus 4 minutes of developer attention, the monthly waste is brutal:
- 60 flaky jobs/day
- 60 x 5.5 minutes of combined time = 330 minutes/day
- Roughly 165 hours/month of human time
- About $1,200-$3,500/month in extra CI compute, depending on runner class and cloud pricing
That is not a rounding error. It is a tax on every release.
Why "pass on retry" is a false green
A retry masks the symptom while preserving the cause. If a test passes only after the fifth run, you do not have a stable test with bad luck. You have a nondeterministic dependency, an order-sensitive assertion, a race condition, or an environment mismatch.
A typical pattern looks like this:
# Anti-pattern: retrying until the signal looks green
jobs:
test:
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@v4
- run: npm ci
- run: npm test -- --retry=5
That pipeline reports success, but it also erases the evidence you need to fix the root cause. In 2026, with larger monorepos and heavier parallelization, that kind of masking is more expensive than ever.
Where flaky tests come from in modern CI systems
Flaky tests usually come from one of five sources: timing, shared state, external dependencies, environment drift, and concurrency.
1. Timing and async race conditions
A test that waits 500 ms for an event that arrives in 650 ms will fail intermittently forever. This is common in UI tests, event-driven services, and streaming pipelines.
Example: a checkout service emits OrderReserved after a Kafka consumer ack. On a quiet runner, the event arrives in 120 ms. Under load, it arrives in 1.8 seconds. A test with a hardcoded 1-second sleep fails 1 in 8 runs.
Use condition-based waits instead of fixed sleeps:
// Better: wait for the state you actually need
await waitFor(async () => {
const order = await db.orders.findById(orderId);
expect(order.status).toBe('RESERVED');
}, { timeout: 5000, interval: 100 });
2. Shared state and test pollution
If one test writes to a global cache, temp directory, Redis key, or feature flag and another test reads it, your suite becomes order-dependent. That is how a test passes locally and fails in CI only when shard 7 runs before shard 8.
A real-world symptom: a Node.js test suite with 1,100 specs passed at 98.9% locally but dropped to 93.4% in parallel CI because three tests reused the same tenant_id=demo fixture.
3. External dependencies with unstable contracts
Payment gateways, identity providers, and third-party APIs can be stable and still make your tests flaky. Rate limits, sandbox resets, expired tokens, and delayed webhooks create false failures.
The fix is not "mock everything." The fix is to isolate the boundary. Use contract tests for provider behavior, and reserve a small number of end-to-end tests for true integration checks.
4. Environment drift
A test that depends on local time zone, CPU count, filesystem semantics, or browser version is fragile by design. In 2026, this is more common because teams run mixed fleets: x86 runners, ARM runners, containerized browsers, and ephemeral preview environments.
A classic example is date parsing:
# Flaky across time zones and locales
assert parse_date("03/04/2026") == date(2026, 3, 4)
If the runtime locale changes, the test may flip meaning. Use explicit formats and UTC where possible.
5. Concurrency and resource contention
Parallel test execution exposes hidden assumptions. Two tests hitting the same port, file, or queue can collide. In large CI systems, this often appears only after scaling from 4 to 16 runners.
One platform team at a fintech reduced flake rate from 3.1% to 0.4% after they stopped sharing a single PostgreSQL schema across shards and started provisioning per-shard schemas in under 11 seconds each.
How flaky tests distort engineering decisions
Flaky tests do more than slow you down. They change how teams ship.
They lower trust in the pipeline
Once engineers see repeated false alarms, they stop treating red builds as urgent. Mean time to acknowledge rises. In one SaaS org, the average response time to a failed main-branch build grew from 9 minutes to 41 minutes after flake rate crossed 2%.
They inflate the cost of change
A noisy suite makes every refactor feel risky. Teams avoid touching brittle code because they expect a cascade of failures. That leads to architectural stagnation: old code stays untouched because the test suite cannot be trusted.
They hide real regressions
The most dangerous outcome is habituation. If a test fails five times and passes once, the one pass becomes a comforting lie. When a real defect appears in the same area, it blends into the noise.
A flaky test is not a low-priority defect. It is a degraded control system.
A practical system for finding and fixing the fifth-run failure
You do not fix flakiness by adding more retries. You fix it by making nondeterminism visible, then removing it.
Step 1: Measure flake rate by test, not by pipeline
Track failures per test case, per branch, per runner type, and per time window. A single flaky test with a 12% failure rate matters more than a suite-wide average of 0.8%.
A useful threshold model:
- 0-0.5%: acceptable, monitor weekly
- 0.5-2%: investigate, prioritize by runtime and business criticality
-
2%: block merges until fixed or quarantined with an owner
Step 2: Capture the right evidence
When a test fails, store:
- seed values
- environment variables
- container image digest
- browser version
- shard ID
- timestamps with millisecond precision
- network traces or HTTP logs for integration tests
A minimal example for Python test runs:
pytest tests/ -q \
--maxfail=1 \
--durations=20 \
--json-report \
--json-report-file=artifacts/pytest-report.json
Step 3: Remove randomness from the test path
If you use randomized IDs, clock-based logic, or unordered collections, make them deterministic in test mode.
// Make time injectable instead of reading the system clock directly
type Clock interface {
Now() time.Time
}
type FixedClock struct{ t time.Time }
func (c FixedClock) Now() time.Time { return c.t }
Step 4: Isolate infrastructure from behavior
Use test containers, ephemeral databases, and per-test namespaces. If a test depends on Redis, give it its own keyspace. If it depends on a queue, give it a unique topic or prefix.
Architecture pattern:
CI Runner
-> ephemeral namespace per shard
-> dedicated Postgres schema per shard
-> unique Redis prefix per test suite
-> mocked third-party APIs at the edge
-> small set of contract tests against real providers
Step 5: Quarantine with expiration, not amnesia
If you must quarantine a flaky test, attach an owner and an expiry date. A quarantine without a deadline becomes permanent technical debt.
A good policy is 7 days for high-severity suites and 14 days for noncritical coverage, with automatic escalation if the test is still quarantined.
Common Pitfalls
"Just add retries"
Retries are useful for proving flakiness, not for hiding it. If a test only passes on the fifth run, retries have turned a signal into noise.
Over-mocking everything
Mocking external systems can make tests fast but unrealistic. If your contract with Stripe, Auth0, or a partner API matters, keep one contract layer against the real schema.
Sharing fixtures across tests
Global fixtures save time and destroy isolation. Use factory functions and unique identifiers per test run.
Ignoring runner differences
A suite that passes on Linux x86 but flakes on ARM64 or on macOS preview runners has an environment bug, not a "CI issue."
Quarantining without ownership
If nobody owns the flake, nobody fixes it. Assign a team, a deadline, and a rollback path.
What a stable pipeline looks like in 2026
A modern 2026 pipeline is not just faster; it is more diagnosable. Teams are combining ephemeral environments, test observability, and policy-based gating to keep flake rates below 1%.
A realistic target for an enterprise platform team:
- Median test runtime: under 18 minutes for full suite
- Flake rate: under 0.7% per 1,000 runs
- Mean time to isolate root cause: under 30 minutes
- Retry usage: limited to infrastructure failures, not assertion failures
One engineering org cut false red builds by 72% after moving from shared staging to per-branch preview environments and tagging every test with a deterministic seed. Their CI spend rose 8% from more ephemeral environments, but developer time saved offset that within six weeks.
Key Takeaways
- Treat a test that passes only on the fifth run as a defect, not a nuisance.
- Measure flake rate per test case and block merges when critical tests exceed your threshold.
- Replace fixed sleeps with condition-based waits and deterministic clocks.
- Isolate databases, queues, caches, and namespaces per test or per shard.
- Use retries only to confirm flakiness, then remove the cause.
- Quarantine failing tests with an owner and expiry date so debt does not become permanent.
- Build CI around observability: seeds, logs, shard IDs, image digests, and timestamps.
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Written by
Nesqual Tech AI
Nesqual Tech
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI