Retries Won’t Fix a Bad Locator 

A login test passes on Monday, fails on Tuesday, and passes again on Wednesday. Nobody touched the login flow. What changed was a container the design team added around the sign-in button, which moved the button one level deeper in the screen’s element tree. The XPath pointing at that button was written against the old layout. Depending on how fast the screen rendered, the lookup sometimes landed and sometimes didn’t. 

That’s what flakiness often looks like in an Appium suite. Not a mystery, but a test that depends on details of the UI it was never meant to check. 

Flakiness has more than one cause 

A flaky test is one that passes and fails on the same code. At scale, it’s expensive. Atlassian reports that flaky tests account for roughly 15% of failures in its Jira backend repository, driving reruns that waste over 150,000 developer hours a year, and were responsible for up to 21% of main-branch build failures in Jira Frontend (Atlassian, 2025). 

The causes sort into two buckets: 

  • Test-owned causes. Fragile locators, fixed sleeps instead of real waits, shared test data, order dependencies. These live in your code, and you can fix them. 
  • Environment causes. A device that drops off the network, a slow backend call, an unexpected OS dialog. These live outside your code. You can reduce them, but not eliminate them. 

The distinction matters because the fixes are different. Retrying a test with a broken locator just produces the same failure more slowly, or worse, an occasional pass that hides the problem. Start with what you own. In mobile UI tests, locators are usually the first place to look. 

Why locators break when the app didn’t 

Every Appium step starts with finding an element. The locator is the address the test uses to find it. When that address describes the structure of the screen rather than the identity of the element, any layout change can break it. 

Researchers call this locator fragility. A 2026 study analyzed 359 open-source repositories and reproduced 449 locator breaks: cases where a structural change to the UI caused a test to fail even though the application still behaved correctly (ReproBreak, ASE 2026). That research covered web test frameworks, but the mechanism is the same on mobile. Rename an element, add a container, reorder a list, and a structure-based locator points somewhere else. 

Mobile adds two more pressures: 

  • Platform differences. The same screen produces a different element tree on iOS and Android, so one XPath rarely works on both. 
  • Rendering timing. Mobile screens load asynchronously. A correct locator can still fail if the test looks before the element exists, which is why locator and wait problems often show up together. 

AI-generated tests raise the stakes 

Test volume is growing faster than teams can review it. An estimated 40–50% of enterprise code is now AI-generated (Digital.ai), and a growing share of test code is generated too. 

Slack’s engineering team recently ran AI-generated Playwright tests against real workflows. The generated tests failed about 8% of the time on a simple flow and about 48% of the time on a more complex one. Slack attributed the failures primarily to variability in UI state and to existing abstractions that interfered with precise element targeting (Slack Engineering, 2026). That was web, not mobile, but the lesson carries: generated tests inherit whatever locator habits they’re given. If no one reviews how elements are found, a bigger suite means more fragile lookups. 

A practical locator order for Appium 

The Appium project’s own guidance ranks locator strategies in this order, and calls XPath the last resort (Appium): 

 

  1. Accessibility ID first. It’s set deliberately by the app team, it’s fast to look up, and it uses the same strategy on iOS and Android. 
  2. Resource ID or ID next. Stable as long as developers don’t rename them. 
  3. Platform-native queries when no ID exists. iOS predicate strings and class chains, or Android UiAutomator. More precise than XPath, but platform-specific. 
  4. XPath last. Appium’s guidance describes it as slow and brittle. When you must use it, keep expressions short and attribute-based rather than index-based. 

Habits that keep locators stable 

  • Make testability part of the definition of done. Ask app teams to add accessibility identifiers to interactive elements. It also improves accessibility for real users. 
  • Centralize locators. Keep them in page objects or screen classes, so a UI change means one edit, not twenty. 
  • Wait for conditions, not time. Replace fixed sleeps with explicit waits for an element to be present or clickable. 
  • Review locators in code review, including generated tests. An XPath with three indexes in it is a future failure. 
  • Track lookup failures separately. If element-not-found errors cluster around a few screens, that’s where locator cleanup pays off first. 

Fix what you own, then manage what you don’t 

Stable locators remove a large class of failures that are fully in your control. They won’t remove the ones that aren’t. Devices, networks, and environments still produce intermittent failures, and then the question becomes how quickly your team can tell those apart from real defects without rerunning tests by hand. 

That’s the second half of the flakiness problem, and it’s the subject of our next post: Retries Are Easy. Knowing What They Told You Isn’t 

You Might Also Like