Functional, including the unhappy paths.
Every flow walked to its end, including the checkout with each instrument the market actually uses and the failure case behind each one
a case list somebody can rerun rather than a session somebody remembers
App testing and QA · UAE
A build can pass a full cycle in English on a recent handset and fail its first Arabic customer. The defect this market produces is a locale defect, and it never reproduces in English.
Find out what the release is carryingWhat app testing and QA is
Three things decide whether a pass means anything here. Only one of them is the device, and it is the one every quote already covers.
An Arabic build is a layout condition rather than a translation. The same screen mirrors, and the same label runs half again as long, so the second script is an axis on the matrix instead of a checkbox at the end.
The checkout has more than one instrument behind it, and a postcode does not locate an address here the way it does where the form was designed. A checkout test that passed on a card form has tested one path through a screen with several.
A pass or a fail against a named standard, with reproducible steps. What is not fixed is stated rather than omitted, because a release decision made on a list of opinions is not a decision.
The scope
Six passes over the same build. Two of them find things the other four are structurally unable to see.
Every flow walked to its end, including the checkout with each instrument the market actually uses and the failure case behind each one
a case list somebody can rerun rather than a session somebody remembers
Cold start, memory and the screens that stutter, recorded as numbers against a budget agreed beforehand
a figure to argue with instead of an opinion about whether it feels slow
The OWASP mobile checklist worked through item by item, with each finding rated rather than listed
a severity-ordered set you can decide about, and a written note of what was accepted
The matrix run across handsets and operating system versions, and then again in the other reading direction
the grid where a locale defect becomes visible, because it has nowhere else to appear
The regression-prone paths scripted and run on every build, in your own pipeline rather than on our machine
a suite that stays yours when the engagement ends
A person using the app badly on purpose, then filing what broke with the steps, the device and the build number
bugs a developer can reproduce on the first attempt
A gate, or the reviews
Every build gets tested. The only question is whether it happens before the release or in the store listing afterwards.
| Ship and find outFound in the reviews | A gate before releaseFound in the matrix | |
|---|---|---|
| Who finds it | A customer, in public, with a star rating attached to the discovery. | A tester, on a Tuesday, with a build number and steps to reproduce. |
| The second script | Never tested, because it looked fine in the screenshots somebody approved. | An axis on the matrix, so a mirrored layout fault has somewhere to appear. |
| Regressions | Found by whoever happens to use that screen next, weeks later. | Caught on the build that caused them, by a suite that runs unattended. |
| What you get | A support queue, and a rating that takes months of good releases to move. | A pass or a fail, and a written list of what is knowingly not fixed. |
Enough to make the release decision an informed one, which is less than a full regression on every build and a good deal more than opening the app once. The shape that works is a small automated suite on the flows that keep breaking, plus a human pass in both scripts before anything ships. Where nobody has touched the app in years the first question is a different one, answered on the maintenance page, and the device report this produces is also a build deliverable, covered on the Dubai page.
What we find in testing here
Webzenia has worked with Gulf clients since 2018. Four things that decide whether a pass is worth anything in this market.
Android is 77.53% of use and iOS 22.46%, and the two most recent Android releases carry roughly half of Android use while older ones still hold a real tail. That is a smaller compatibility bill than most matrices assume, and the room it frees is where the second script gets tested.
A label that fits in English runs half again as long, a mirrored screen moves the back gesture, and an icon that meant forward now points away from it. A Business Bay proptech on a mainland DET licence whose Arabic came from a translation vendor has never had any of it opened by a tester.
A wallet, a card, an instalment plan and cash on delivery are four routes through one screen, each failing differently, and the address form is a fifth because a postcode does not locate a delivery point. An Al Quoz F&B group on a mainland DET licence found that in its store rating.
A library added late brings its own defaults, and nobody updates the declaration. Apple has required a privacy manifest from newly added third-party SDKs since 1 May 2024, so the gate is whether the Tabby, Tamara or Network International wrapper declares itself. A Sharjah SAIF Zone manufacturer finds the gap at submission.
What the engagement covers
Six stages, and the first one is a conversation about risk rather than a list of screens. What can go wrong expensively decides what gets tested first.
What fails expensively, on which handsets and in which scripts, from your own analytics rather than from a generic device list. The plan is what makes the coverage arguable rather than assumed.
Written so a person who has never seen the app can execute them and get the same answer. A case only its author can run is a memory rather than a test.
The regression-prone flows scripted into your own pipeline, so they run on every build without anybody deciding to. Automating everything is a project; automating the fragile paths is a habit.
The matrix executed on physical devices and then executed again in the other reading direction, because a mirrored layout is a different render and not a different string.
The OWASP mobile items worked through with each finding rated, plus the library inventory taken from the shipped binary and compared with what the listing declares.
One decision, with the severity list behind it and the accepted items written down. A report that ends in a bug count leaves the release decision to whoever is most tired.
Our stack
One suite, a real device matrix, a way to break the network, a crash log and the repository it all lives in. Select one to see what it catches.
One suite that drives both platforms means a regression case is written once, and a locale run is the same suite with the device language changed rather than a second set of scripts.
We parameterise the suite by locale as well as by device, so every automated case runs in both reading directions. A suite that only runs in English automates the half of the problem that was never failing.
How the work runs
Testing is worth what the decision at the end of it is worth. The output is a release call, not a bug count.
We establish what fails expensively, which handsets and operating system versions your own users are actually on, and which locales the product ships in. The cases are then written so somebody who has never seen the app can execute them. This week decides the coverage, and coverage decided by a device catalogue tests somebody else’s users thoroughly.
The suite runs on physical handsets and then again in the other reading direction, because a mirrored layout is a different render rather than a different string. Each payment instrument is walked to completion and to failure, the network is broken deliberately, and the regression-prone flows are automated into your own pipeline as they are found.
Each release gets a decision rather than a document: pass, or fail with the reason. What is knowingly not fixed is written down and accepted by somebody named, so a defect that ships is a choice rather than an oversight. Where the release rhythm is weekly this becomes a standing arrangement, which is a different commercial shape from a one-off pass and usually the right one.
Reported as a pass or a fail, so what is not fixed is a decision somebody made.
Our commitment
Testing goes wrong when it produces a document instead of a decision. These four are written into the scope for that reason.
We test Arabic, or say we are not.
Every case runs in both reading directions, or the engagement states plainly that the Arabic build was not covered. What we will not do is test around it and hand back a pass that means less than it appears to.
The phone list comes from your own analytics.
Which handsets and operating system versions we run on is decided by what your users are actually holding. A standard device list is somebody else’s user base, tested carefully.
A pass or a fail, with gaps named.
One decision, the severity list behind it, and the accepted items written down. A report that ends in a bug count hands the release call back to whoever is most tired at the end of the sprint.
The test suite is yours.
Cases and automation live in your repository and run in your pipeline from the first week. A gate that only we can run is not a gate you own, whatever the retainer calls it.
Common questions
Keep exploring
Native Swift, and the entity Apple publishes as seller.
Kotlin, and the five manufacturer skins this market runs.
One codebase, and the parts written per platform.
One widget tree, and an Arabic screen that mirrors itself.
For teams that already write React and have to hold it.
A shopping app costed against the repeat order.
Two-sided products where the fleet is the harder half.
Internal apps for staff who are not at a desk.
Subscription products whose customer is a licensed company.
Next step
Send us the build and the two or three flows that matter most. We will come back with what fails, on the handsets this market carries and in both scripts.
Tell us what you need.