⚙ Test Automation

Why your test suite is flaky — and 5 fixes that work

Flaky tests are not bad luck. They are a category of bug with known causes — and known cures. Here is how we diagnose them and the rewrite patterns that make a suite trustworthy again.

Every engineering team eventually encounters a flaky test — a test that passes, fails, and then passes again without a relevant code change. It may pass locally, fail in CI, and then pass again when re-run. Someone clicks "re-run," it goes green, and everyone moves on. That habit is dangerous because once a team stops trusting a red build, the entire test suite stops providing reliable feedback.

The good news: flakiness is not mysterious. After enough engagements, the same handful of root causes account for nearly every flaky test we have ever triaged. Fix those patterns and a suite goes from "ignore the failures" to "red means stop" — which is the only state in which automation is worth running at all.

01❗Why Flaky Tests Are a Problem

A single flaky test seems harmless. The damage is cultural, not technical. The first time a real failure gets dismissed as "probably just flaky" and shipped anyway, the suite has inverted its purpose: it is now training your team to ignore the signal it exists to send. Velocity drops too — every re-run is wasted CI time and a context switch — but trust is the real loss, and trust does not come back with a re-run.

✓

The rule we enforce: a flaky test is quarantined within 24 hours — pulled from the blocking suite, root-caused, then fixed or deleted. A test nobody trusts is worse than no test, because it carries the cost without the protection.

02❗5 Common Causes of Flaky Tests

Almost every flake we have diagnosed traces back to one of these. None of them require luck to fix — each has a specific, repeatable cure.

  • Timing assumptions — the test waits a fixed number of seconds and hopes the app is ready. Sometimes it is not.
  • Test interdependence — test B only passes if test A ran first, so any reorder or parallel run breaks it.
  • Shared mutable state — tests fight over the same database row, user account, or fixture.
  • Non-deterministic data — today's date, random IDs, or unsorted API responses leak into assertions.
  • Real network calls — a third-party API is slow or down, and your "test failure" is actually their outage.

Common Cause

What Happens

Typical Solution

Timing assumptions

The application is not ready when the test runs

Wait for the required condition

Test interdependence

One test depends on another test

Keep tests independent

Shared mutable state

Tests modify or compete for the same data

Use isolated test data

Non-deterministic data

Dates, random values, or ordering change results

Control test data and randomness

Real network calls

External services cause unpredictable failures

Mock external dependencies

03❗Fix #1: Wait for Conditions, Not the Clock

The single biggest source of flakiness is the hard-coded wait. The test pauses for two seconds because that was "usually enough." On a slow CI runner, it is not enough, and the test fails on an element that was about to appear. The fix is to wait for the state you actually care about, not for the clock. Toggle these two and watch the difference:

checkout.spec.ts
await page.click('.submit'); await page.waitForTimeout(2000); // ← hoping 2s is enough const toast = await page.textContent('.toast'); expect(toast).toBe('Saved'); // ← fails when the runner is slow
await page.getByRole('button', { name: 'Submit' }).click(); // waits for the condition, however long it takes (up to a cap) await expect(page.getByText('Saved')).toBeVisible();
✕ passes ~86% of runs — the 14% are slow-runner false alarms that erode trust

The stable version is also shorter and clearer. That is the usual pattern: the flaky-proof way to write a test is almost always the more readable way too.

04❗Fix #2: Isolate Tests and Test Data

Tests that share a user account, a database row,w or a logged-in session will eventually collide — especially once you run them in parallel to go faster. The cure is isolation: each test creates the data it needs and cleans up after itself, so order never matters and parallelism is free. Set up fresh fixtures per test; never rely on a record a previous test left behind.

▲

Watch for the "works alone, fails together" tell. If a test passes when you run it by itself but fails in the full suite, you almost certainly have shared state — not a flaky framework.

05❗Fix #3: Control Time, Randomness, and Test Data

If an assertion depends on today's date, a random ID, or the order of an unsorted API response, it will pass until the day it does not. Freeze the clock in tests, seed or stub random generators, and sort collections before asserting on them. Determinism is not a nice-to-have — it is the definition of a test that means something.

06❗Fix #4: Avoid Unnecessary External Dependencies

When a test depends on a real third-party API, your build reliability also depends on a service you do not control. This is especially important when building reliable automation testing workflows, where unstable dependencies can create misleading failures. For everything outside your system boundary, mock it — assert that your code handles the contract correctly. Save the real integration calls for a separate, non-blocking suite whose failures alert you without stopping a deploy.

# the state every suite should reach flake_budget: 0 tolerated quarantine_sla: 24 hours red_means: "stop and look" — every time ✓

07❗Fix #5: Make Flaky-Test Reduction a Team Policy

The four fixes above are techniques. The fifth is the discipline that keeps them in place: track flaky tests as a first-class quality metric and investigate them quickly. Track flakiness as a first-class metric and use appropriate bug tracking tools to record, prioritize, and monitor recurring test failures, and treat a flaky test as a P2 bug in your test suite because that is exactly what it is. Suites do not stay trustworthy by accident; they stay trustworthy because someone made untrustworthiness a thing you are not allowed to ship.

Do this, and your team gets the thing automation was always supposed to deliver: a green build that actually means "safe to ship," and a red build that everyone stops for. That is not a luxury. It is the whole point.

❕How to Identify a Flaky Test

A test may be flaky when it produces different results under substantially the same conditions.

Common signs include:

  • The test passes and fails without a relevant code change.
  • Re-running the same test produces a different result.
  • The test fails more often in CI than locally.
  • The failure appears only during parallel execution.
  • The failure disappears when the test is re-run.
  • The problem cannot be reproduced consistently.
✓

Drowning in flakes already? A suite stabilization pass is one of the fastest-payback engagements we run — we triage, quarantine and rewrite until red means red again. Build a plan or book a scoping call below.

08❗Flaky Test vs Failed Test

Flaky Test

Failed Test

May pass or fail between runs

Usually fails consistently under the same conditions

Often caused by non-deterministic behavior

May indicate a reproducible application or test issue

Failure can disappear on re-run

Re-running usually produces the same result

Requires investigation into test stability

Requires investigation into the underlying failure

A failed test is not automatically a flaky test. A test is considered flaky when its result is inconsistent under substantially the same conditions.

❕How to Prevent Flaky Tests

You can reduce flaky tests by:

  • Using condition-based waits instead of fixed delays.
  • Keeping tests independent.
  • Creating isolated test data.
  • Controlling dates, times, and random values.
  • Avoiding unnecessary external dependencies.
  • Making parallel test execution safe.
  • Tracking flaky-test failures in CI.
  • Investigating intermittent failures instead of repeatedly re-running them.

09❗Frequently Asked Questions

👉 What is a flaky test?

A flaky test is an automated test that produces inconsistent results without a relevant change to the application or test code.

👉 Why do tests pass locally but fail in CI?

Tests can behave differently in CI because of differences in timing, environment, resources, parallel execution, test data, or external dependencies.

👉 What causes flaky tests?

Common causes include timing assumptions, test dependencies, shared state, non-deterministic data, and external network dependencies.

👉 How do you fix a flaky test?

Identify the source of non-deterministic behavior, then isolate test data, replace fixed waits with condition-based waits, control time and randomness, and reduce unnecessary external dependencies.

👉 Should flaky tests be ignored?

No. Ignoring flaky tests reduces confidence in the test suite and can cause genuine failures to be overlooked.

A reliable automated test suite should make failures meaningful. When a build is green, the team should have confidence that the tested functionality is working as expected. When a build is red, the failure should be investigated rather than automatically dismissed as flaky. Reducing flaky tests is therefore essential to maintaining trust in test automation.

TM
Twisa Mistry

Senior QA engineers who have stabilized suites across SaaS, FinTech and Enterprise teams since 2017.

Want red to mean red again?

Bring us your flakiest suite. A stabilization pass is one of the fastest-payback things we do.

Book a Scoping Call