Every engineering team eventually encounters a flaky test — a test that passes, fails, and then passes again without a relevant code change. It may pass locally, fail in CI, and then pass again when re-run. Someone clicks "re-run," it goes green, and everyone moves on. That habit is dangerous because once a team stops trusting a red build, the entire test suite stops providing reliable feedback.
The good news: flakiness is not mysterious. After enough engagements, the same handful of root causes account for nearly every flaky test we have ever triaged. Fix those patterns and a suite goes from "ignore the failures" to "red means stop" — which is the only state in which automation is worth running at all.
01❗Why Flaky Tests Are a Problem
A single flaky test seems harmless. The damage is cultural, not technical. The first time a real failure gets dismissed as "probably just flaky" and shipped anyway, the suite has inverted its purpose: it is now training your team to ignore the signal it exists to send. Velocity drops too — every re-run is wasted CI time and a context switch — but trust is the real loss, and trust does not come back with a re-run.
The rule we enforce: a flaky test is quarantined within 24 hours — pulled from the blocking suite, root-caused, then fixed or deleted. A test nobody trusts is worse than no test, because it carries the cost without the protection.
02❗5 Common Causes of Flaky Tests
Almost every flake we have diagnosed traces back to one of these. None of them require luck to fix — each has a specific, repeatable cure.
- Timing assumptions — the test waits a fixed number of seconds and hopes the app is ready. Sometimes it is not.
- Test interdependence — test B only passes if test A ran first, so any reorder or parallel run breaks it.
- Shared mutable state — tests fight over the same database row, user account, or fixture.
- Non-deterministic data — today's date, random IDs, or unsorted API responses leak into assertions.
- Real network calls — a third-party API is slow or down, and your "test failure" is actually their outage.
|
Common Cause |
What Happens |
Typical Solution |
|---|---|---|
|
Timing assumptions |
The application is not ready when the test runs |
Wait for the required condition |
|
Test interdependence |
One test depends on another test |
Keep tests independent |
|
Shared mutable state |
Tests modify or compete for the same data |
Use isolated test data |
|
Non-deterministic data |
Dates, random values, or ordering change results |
Control test data and randomness |
|
Real network calls |
External services cause unpredictable failures |
Mock external dependencies |
03❗Fix #1: Wait for Conditions, Not the Clock
The single biggest source of flakiness is the hard-coded wait. The test pauses for two seconds because that was "usually enough." On a slow CI runner, it is not enough, and the test fails on an element that was about to appear. The fix is to wait for the state you actually care about, not for the clock. Toggle these two and watch the difference:
The stable version is also shorter and clearer. That is the usual pattern: the flaky-proof way to write a test is almost always the more readable way too.
04❗Fix #2: Isolate Tests and Test Data
Tests that share a user account, a database row,w or a logged-in session will eventually collide — especially once you run them in parallel to go faster. The cure is isolation: each test creates the data it needs and cleans up after itself, so order never matters and parallelism is free. Set up fresh fixtures per test; never rely on a record a previous test left behind.
Watch for the "works alone, fails together" tell. If a test passes when you run it by itself but fails in the full suite, you almost certainly have shared state — not a flaky framework.
05❗Fix #3: Control Time, Randomness, and Test Data
If an assertion depends on today's date, a random ID, or the order of an unsorted API response, it will pass until the day it does not. Freeze the clock in tests, seed or stub random generators, and sort collections before asserting on them. Determinism is not a nice-to-have — it is the definition of a test that means something.
06❗Fix #4: Avoid Unnecessary External Dependencies
When a test depends on a real third-party API, your build reliability also depends on a service you do not control. This is especially important when building reliable automation testing workflows, where unstable dependencies can create misleading failures. For everything outside your system boundary, mock it — assert that your code handles the contract correctly. Save the real integration calls for a separate, non-blocking suite whose failures alert you without stopping a deploy.
07❗Fix #5: Make Flaky-Test Reduction a Team Policy
The four fixes above are techniques. The fifth is the discipline that keeps them in place: track flaky tests as a first-class quality metric and investigate them quickly. Track flakiness as a first-class metric and use appropriate bug tracking tools to record, prioritize, and monitor recurring test failures, and treat a flaky test as a P2 bug in your test suite because that is exactly what it is. Suites do not stay trustworthy by accident; they stay trustworthy because someone made untrustworthiness a thing you are not allowed to ship.
Do this, and your team gets the thing automation was always supposed to deliver: a green build that actually means "safe to ship," and a red build that everyone stops for. That is not a luxury. It is the whole point.
❕How to Identify a Flaky Test
A test may be flaky when it produces different results under substantially the same conditions.
Common signs include:
- The test passes and fails without a relevant code change.
- Re-running the same test produces a different result.
- The test fails more often in CI than locally.
- The failure appears only during parallel execution.
- The failure disappears when the test is re-run.
- The problem cannot be reproduced consistently.
Drowning in flakes already? A suite stabilization pass is one of the fastest-payback engagements we run — we triage, quarantine and rewrite until red means red again. Build a plan or book a scoping call below.
08❗Flaky Test vs Failed Test
|
Flaky Test |
Failed Test |
|---|---|
|
May pass or fail between runs |
Usually fails consistently under the same conditions |
|
Often caused by non-deterministic behavior |
May indicate a reproducible application or test issue |
|
Failure can disappear on re-run |
Re-running usually produces the same result |
|
Requires investigation into test stability |
Requires investigation into the underlying failure |
A failed test is not automatically a flaky test. A test is considered flaky when its result is inconsistent under substantially the same conditions.
❕How to Prevent Flaky Tests
You can reduce flaky tests by:
- Using condition-based waits instead of fixed delays.
- Keeping tests independent.
- Creating isolated test data.
- Controlling dates, times, and random values.
- Avoiding unnecessary external dependencies.
- Making parallel test execution safe.
- Tracking flaky-test failures in CI.
- Investigating intermittent failures instead of repeatedly re-running them.
09❗Frequently Asked Questions
👉 What is a flaky test?
A flaky test is an automated test that produces inconsistent results without a relevant change to the application or test code.
👉 Why do tests pass locally but fail in CI?
Tests can behave differently in CI because of differences in timing, environment, resources, parallel execution, test data, or external dependencies.
👉 What causes flaky tests?
Common causes include timing assumptions, test dependencies, shared state, non-deterministic data, and external network dependencies.
👉 How do you fix a flaky test?
Identify the source of non-deterministic behavior, then isolate test data, replace fixed waits with condition-based waits, control time and randomness, and reduce unnecessary external dependencies.
👉 Should flaky tests be ignored?
No. Ignoring flaky tests reduces confidence in the test suite and can cause genuine failures to be overlooked.
A reliable automated test suite should make failures meaningful. When a build is green, the team should have confidence that the tested functionality is working as expected. When a build is red, the failure should be investigated rather than automatically dismissed as flaky. Reducing flaky tests is therefore essential to maintaining trust in test automation.
