QA / Software Testing

Flaky Tests: Why Automated Tests Fail Randomly and How to Fix Them

JIN

Oct 09, 2026

Table of contents

Table of contents

    Your pull request is approved; the change is small, and CI turns red on a test you never touched. You hit rerun. It passes. You merge and move on, and nobody asks why it failed, because everybody has seen it before.

    That shrug is how flaky tests do their damage. One flaky failure costs a few minutes. Hundreds of them teach a whole team that a red build is probably noise, and that is the moment a test suite stops protecting you. Google found that 84% of the pass-to-fail transitions in its post-submit CI involved a flaky test. GitHub measured that 1 in 11 commits in its monolith had a red build caused by one.

    This guide covers what a flaky test is, why it matters, the seven causes behind most of them, how to find them, a step-by-step way to fix them, how to manage the ones you cannot fix today, and where AI helps and where it does not. Every figure comes from a named source, listed at the end.

    What Is a Flaky Test?

    A flaky test is an automated test that both passes and fails when run against the same code, with no change to the test or the system under test. The outcome depends on something the test does not control, such as timing, execution order, shared data, or the environment. This is the definition Google’s testing team uses.

    You will also hear the term nondeterministic test, and the symptom described as an intermittent test failure. They all point to the same problem: the result is not a reliable function of the code.

    The classic mistake is a fixed pause before clicking a button that loads asynchronously:

    // Flaky: guesses that the page is ready after 2 seconds
    await page.waitForTimeout(2000);
    await page.click('#submit');

    // Stable: waits for the button itself
    await page.getByRole('button', { name: 'Submit' }).click();

    The two-second sleep is a guess. On a fast laptop, the button is ready in 400 milliseconds, so the test wastes time. On a busy CI runner, the API takes 2.5 seconds, so the click lands before the button is usable. Same code, two outcomes.

    Flaky Test vs. Failing Test vs. Flaky Product

    Not every random red build is a test problem. Sorting the failure into the right bucket decides who fixes it.

    Type How it behaves What it means Right response
    Failing test Fails on every run A real bug, or a broken test Fix the code or the test
    Flaky test Fails on some runs The test or its environment is unreliable Find the root cause in the test, its data, or the environment
    Flaky product Fails on some runs The application has a race condition, and the test is telling the truth File a defect and fix the application

    Why “It Passed on Rerun” Is Not Proof

    A green rerun shows only that the failure is intermittent. It does not show that the failure is harmless. If the cause is a timing defect in the application, the rerun simply produced a different outcome, and the same race can affect a real user in production. Treat a pass-after-fail as a lead to investigate, not as a clean bill of health.

    Why Do Flaky Tests Matter?

    Flaky tests in CI/CD pipelines cost more than the rerun button suggests. The damage shows up in four places.

    1. CI/CD Delays and Blocked Releases

    GitHub reported that in its monolith, 1 in 11 commits (about 9%) had at least one red build caused by a flaky test. Anyone trying to deploy a handful of commits had a good chance of having to retry a build or diagnose a failure, even though their code was fine.

    The math compounds quickly. Suppose each of your 1,000 tests has a 0.1% chance of failing for no reason. The chance that all of them pass is about 37%, so roughly 63% of full runs go red. That is simple arithmetic that assumes independent flakes, not a measured figure, but it shows why a suite that looks mostly stable can feel broken. Every rerun also burns pipeline minutes and holds up the merge queue for everyone behind you.

    2. False Failures and Wasted Engineering Hours

    Atlassian attributes about 15% of Jira backend repository failures to flaky tests and estimates that the reruns waste over 150,000 developer hours a year. The human loop is familiar: investigate, rerun, ping a teammate, rerun again, then give up and merge. None of it produces a feature or fixes a bug.

    3. Lost Trust in the Test Suite

    Google found that 84% of the transitions from pass to fail in its post-submit system involved a flaky test. Put plainly, most red builds were noise. When that happens, people stop reading red builds, and the one that signals a real regression gets rerun along with the rest.

    4. Hidden Product Bugs

    Some flakes are genuine timing or concurrency defects in the application. A blanket retry policy turns those into green builds and ships them to production. The test was right, and the pipeline taught everyone to ignore it.

    The real cost is not the rerun. It is the gradual decline in decision quality. A pipeline that no longer separates signal from noise is just an expensive delay.

    What Causes Flaky Tests?

    The most cited breakdown comes from Luo et al. (FSE 2014), who studied 201 flaky test fixes across Apache projects. Async waits accounted for 45% of them, concurrency for 20%, and test order dependency for 12%. The study predates modern browser automation, so UI selector drift is probably underrepresented in those numbers. The seven causes below cover what we see most often in practice.

    1. Timing and Async Waits

    This is the largest single cause. A test clicks before the element exists, asserts before an API response arrives, or waits for a page load event when the control it needs renders a second later. Animations, debounced inputs, and background polling jobs open the same kind of gap. The fix is to wait for a condition, not a duration (see step 3 in the fix section below).

    2. Test Order and Hidden Dependencies

    Test B quietly relies on data or setup created by test A. Run B alone, in parallel, or in a shuffled order, and it fails. The same family includes unpinned third-party libraries and framework upgrades that change the order in which tests execute, so a suite that was stable last week breaks after a routine dependency bump. Independent tests and pinned versions remove most of it.

    3. Unstable Test Data

    Shared test accounts, leftover records from earlier runs, hardcoded IDs, expiring dates, and random values with no fixed seed all give a test a different starting point each time. A promo code valid until 31 December works perfectly until January. Create the data each test needs, clean it up afterward, and seed anything random.

    4. Environment Differences

    A test that passes on a laptop and fails on the CI runner is usually seeing a different world: fewer CPU cores, tighter memory, another browser version, a smaller screen, a different timezone or locale, or an older container image. A layout assertion that holds at 1920 pixels wide can fail at 1280. Pin the environment and run tests from the same container image locally and in CI.

    5. Network and External Services

    Tests that call real third-party APIs inherit every hiccup those services have: DNS delays, rate limits, slow staging responses, an outage that has nothing to do with your code. Stub or mock the external call in most tests, and cover the real integration separately with a small set of contract tests.

    6. Shared State and Concurrency

    Global variables, static fields, singletons, caches, and shared databases carry state from one test into the next. When parallel workers write to the same resource, the result depends on who gets there first. One quiet variant is asserting on an unordered collection, such as a hash set or an unsorted query result, as if it had an order. Reset state between tests and give each worker its own resources.

    7. Poor Selectors

    Locators tied to auto-generated CSS classes, deep XPath chains, or an element’s position on the page break when the DOM shifts slightly or renders in a different order. A front-end build that renames classes can turn a whole UI suite red overnight. Prefer data-testid attributes or role-based locators that describe what the element is, not where it sits.

    Cause Symptom First fix to try
    Timing and async waits Passes locally, fails on slow runners; “element not found” Replace sleeps with condition-based waits
    Test order and hidden dependencies Fails when run alone, in parallel, or shuffled Make each test set up its own state
    Unstable test data Fails on a certain date or after earlier runs Create and clean up data per test; seed randomness
    Environment differences Green on a laptop, red in CI Pin versions and use the same container image
    Network and external services Random timeouts and rate-limit errors Mock external calls; add contract tests
    Shared state and concurrency Different results with parallel workers Isolate state per worker; reset between tests
    Poor selectors Breaks after a small UI change Use data-testid or role-based locators

    Most causes come down to one thing: the test assumes something about time, order, or state that nobody guaranteed. Causes also tend to cluster, so one broken fixture or one shared account can make dozens of tests flaky at once. If many tests start failing together, look for the common ingredient before you fix them one by one.

    How Do You Identify Flaky Tests?

    Flaky test detection starts with data, not with gut feeling. Three methods work well together.

    1. Track Failure History per Test

    Record the pass or fail result of every test on every commit. A test that has both results on the same commit is flaky by definition. Then rank tests by how many pull requests or developers they affected, not by raw failure count. GitHub found that most of its flaky tests failed fewer than ten times, and only 0.4% failed 100 times or more. An impact ranking shows you where the real pain is.

    2. Rerun Tests Under Controlled Conditions

    GitHub’s approach is a good worked example. When a test fails, rerun it three ways: in the same process, in the same process with the clock shifted into the future, and on a different host. Each outcome hints at a cause.

    • A pass in the same process points to randomness or a race condition.
    • A pass only with the shifted clock points to a wrong assumption about time.
    • A pass only on a different host points to test order or shared state.

    GitHub says this approach raised automatic detection from 25% to 90% of flaky failures.

    3. Analyze CI Data at Scale

    Once you have history, look for patterns across it. The flip rate of a test, meaning how often its result changes between consecutive runs, is a simple and useful signal. Grouping failures by signature, such as the error message and stack trace, shows when one root cause is behind many tests. Your tools may already help: Playwright reports a test that fails and then passes on retry as flaky, and plugins like pytest-rerunfailures rerun failing tests in Python suites.

    Track three metrics from the start: the flaky test rate, the share of build failures caused by flaky tests, and the time it takes to fix a flaky test once you find it.

    How Do You Fix Flaky Tests?

    Here is how to fix flaky tests in six steps. Follow them in order, because each one protects you from a mistake in the next.

    1. Confirm it is flaky. Reproduce the failure and rerun the test many times before you touch any code. Many teams use 50 to 100 runs as a working bar.

    2. Classify the root cause. Use the failure logs and the rerun pattern to match the test to one of the seven causes above.

    3. Apply the targeted fix. Match the fix to the cause:

    • Timing: use condition-based waits or the framework’s auto-waiting. Playwright and Cypress wait and retry on their own for most actions and assertions, and Selenium needs explicit waits.
    • Data: set up and tear down data per test.
    • Shared state: reset it between tests and isolate it per worker.
    • External calls: mock or stub them, and cover the real integration with contract tests.
    • Environment: pin versions and containerize the CI environment.
    • Time: freeze the clock.
    • Selectors: use stable locators such as data-testid or role-based locators.

    4. Rule out a product bug. If the race is in the application, file it as a defect. Patching the test would hide a real problem.

    5. Verify the fix. Repeat the bulk reruns in CI and under parallel execution, not only on your machine.

    6. Add a guardrail. Add a lint rule that blocks fixed sleeps, a review checklist item, and a test independence check so the same flake does not come back.

    Fixes That Make Things Worse

    • Raising timeouts. A longer timeout only moves the failure to a slower day and makes every run slower.
    • Adding sleeps. A sleep is a guess about timing, and the guess is wrong somewhere.
    • Blanket auto-retry with no tracking. The build turns green, and the problem becomes invisible.
    • Deleting the test without replacing its coverage. The flake is gone, and so is the protection it was supposed to give.

    What Does a Good Flaky Test Management Strategy Look Like?

    You will never reach zero flaky tests. Google reports that its rate of new flaky tests has been about equal to its rate of fixes, and GitHub says it set out to manage the inevitability of flakes, not to eliminate them. So manage them like any other technical debt, with an owner and a deadline. This matters most for flaky tests in CI/CD, where one unreliable test can block everyone.

    1. Detect and Quarantine Without Losing Coverage

    To quarantine flaky tests, keep running them but stop letting them block merges. That keeps the pipeline moving and keeps the data coming. The risk is that quarantine is also a gap in your regression coverage. If nothing pulls tests back out, the quarantine folder becomes a graveyard.

    2. Assign Clear Ownership

    A flaky test with no owner stays flaky. GitHub automatically assigns high-impact flaky tests to the people who most recently changed the test or the related code, and lets teams view them by CODEOWNER. You can copy the idea with CODEOWNERS-style routing so every flaky test lands in a specific team’s queue.

    3. Prioritize by Impact

    GitHub scores each flaky test by the number of failures and by how many branches, developers, and deploys it affected. Because only 0.4% of its flaky tests failed 100 times or more, fixing the top of that list removes most of the pain. Do not try to clean everything at once.

    4. Set a Retry and Quarantine Policy

    Retries are data, not a pass. Log every one. Google, for example, lets teams mark a test as flaky so that it reports a failure only after three consecutive failures, and the mark stays visible. Also set a maximum time a test may sit in quarantine. Per SHIFT ASIA guidance (not cited data), two to four weeks is a sensible starting point, after which the test is fixed or removed on purpose.

    5. Measure and Report

    Track the flaky rate, flaky-caused build failures, quarantine size and age, and mean time to fix. Review them in the same forum as your defect metrics, so test health gets the same attention as product health.

    In Japan-standard QA thinking, a test suite is a quality asset with owners, standards, and audits, not a folder of scripts. Flaky test management is what that idea looks like in daily automation work.

    Can AI Help Detect and Fix Flaky Tests?

    AI is useful for detection and triage today, partly useful for repair, and not a replacement for engineering judgment. AI flaky test detection works best by pointing humans to the right tests faster.

    Where AI Helps Today

    Models can predict which failures are likely flaky from run history, cluster similar failures so you see one root cause instead of fifty tickets, summarize long logs with LLMs, and suggest locator updates when a UI changes.

    What the Research Says About AI Repair

    FlakyDoctor, presented at ISSTA 2024, combines an LLM with program analysis. On 873 confirmed flaky tests from 243 real-world projects, it repaired 57% of the order-dependent ones and 59% of the implementation-dependent ones. Its authors note that LLMs alone were not good enough, since the non-LLM parts contributed 12% to 31% of the result.

    A separate 2024 ICSE student research paper reported higher success on order-dependent and implementation-dependent tests, 79% and 58%, but found that LLMs were ineffective at repairing non-order-dependent flaky tests. That category includes many timing and concurrency problems, the biggest source of flakes in practice. The datasets and methods differ, so read these numbers as a range, not a promise.

    Where AI Falls Short

    Self-healing locators can hide a real UI regression by quietly adapting to a change that users would notice. AI cannot decide whether a flake is a product bug. Every AI-proposed fix still needs repeated reruns and human review, just like a fix written by a person.

    At SHIFT ASIA, we are building our own AI-driven testing framework that covers the work from test case design through execution. Flaky tests are one of the problems tooling should help with, and we will keep the same rule: nothing is trusted until it survives repeated runs.

    Conclusion: Make a Red Build Mean Something Again

    Three things matter most. First, treat every flaky failure as information: a timing assumption, a hidden dependency, or sometimes a real product bug. Second, fix by cause, not by habit, because longer timeouts and blanket retries only hide the problem. Third, treat flaky tests as technical debt: assign owners, score impact, and limit how long a test can stay in quarantine.

    A test suite earns its place only when a red build means something. If your team is rerunning more than it is fixing, SHIFT ASIA’s QA engineers can review your automation suite, find the patterns behind the flakes, and help you build a plan to remove them. Contact us to start the conversation.


    Frequently Asked Questions

     

    Async waits and timing problems are the main causes of flaky tests. In the 2014 study by Luo et al. of 201 flaky test fixes in Apache projects, 45% involved async waits, ahead of concurrency at 20% and test order dependency at 12%. The study predates modern browser automation, so UI selector problems are probably underrepresented.

    Usually the problem sits in the test or its environment, but not always. Some flaky failures come from a real race condition in the application, and the test is reporting it correctly. Check the failure logs and reproduce the issue before you retry or rewrite anything. If the race is in the product, file it as a defect instead of patching the test.

    Only as a short-term safety net. Automatic retries keep pipelines moving, but they also hide the problem, so every retry should be logged and every retried test tracked toward a real fix. A retry that passes tells you the failure is intermittent, not that it is harmless. Without tracking, retries become permanent debt.

    There is no fixed standard, but many teams rerun a suspect test 50 to 100 times. Rare flakes that fail once in a few hundred runs need more. Run the repeats under CI conditions, including parallel workers and the same container image, because a test that is stable on your laptop can still fail on the runner.

    No universal benchmark exists, so set a threshold for your own team and watch the trend. Treat any test that blocks merges as high priority, whatever the overall number. For reference only, Google has reported a continual rate of about 1.5% of test runs showing a flaky result, at a scale far larger than most teams.

    Use condition-based or built-in auto-waiting, stable selectors such as data-testid or roles, isolated test data, and mocked external calls. Playwright's retry report labels a test as flaky when it fails first and passes on retry, which gives you a ready list of suspects. Selenium needs explicit waits, while Cypress retries assertions on its own.

    Partly. Research tools such as FlakyDoctor repaired 57% of order-dependent and 59% of implementation-dependent flaky tests, but results drop for timing-related flakes, and every proposed fix still needs repeated reruns and human review. Today AI is more dependable for detection and triage than for repair.

    Share this article

    ContactContact

    Stay in touch with Us

    What our Clients are saying

    • We asked Shift Asia for a skillful Ruby resource to work with our team in a big and long-term project in Fintech. And we're happy with provided resource on technical skill, performance, communication, and attitude. Beside that, the customer service is also a good point that should be mentioned.

      FPT Software

    • Quick turnaround, SHIFT ASIA supplied us with the resources and solutions needed to develop a feature for a file management functionality. Also, great partnership as they accommodated our requirements on the testing as well to make sure we have zero defect before launching it.

      Jienie Lab ASIA

    • Their comprehensive test cases and efficient system updates impressed us the most. Security concerns were solved, system update and quality assurance service improved the platform and its performance.

      XENON HOLDINGS