Field notes

Designing an End-to-End Test Matrix for a Global Mobile Top-Up Checkout

A worked example of modelling destination, operator, product, payment and fulfilment coverage for a cross-border mobile top-up checkout — using a public flow as the subject.

Testml Desk 8 min read
Isometric diorama of a modular grid of SIM-card-shaped tiles linked by ribbon cables to a central hub, with translucent glass panels above showing node connections.

Cross-border mobile top-up looks like a single button but behaves like a settlement system. A visitor picks a country, types a phone number, chooses an amount, pays with a card or a wallet, and expects an SMS from a mobile operator on the other side of the world within minutes. Each step reaches a different system — carrier catalogue, HLR lookup, card acquirer, 3-D Secure, aggregator — and each system has its own failure modes. Testing the user flow by hand gives you confidence for exactly the handful of combinations you clicked through. Everything else is a trust fall.

This note walks through how we would model coverage for such a product, using the publicly available MobileTopUP checkout as a worked example. The flow is a convenient public reference: a visitor sees it the same way we do when we open the site, and the sequence it exposes (destination → number → operator → product → payment → confirmation) maps cleanly onto a general top-up architecture. Nothing here is a client engagement, an audit or an inside view of that product; MobileTopUP is a shape to think against, not a system we have tested. The point is the matrix, not the merchant.

The shape of the flow

Open the public prepaid top-up flow and the surface decisions are visible on one page: pick a country, enter a mobile number, let the form resolve the operator, choose from the list of top-up products that operator sells, review the price in your local currency, and pay. Behind each visible decision sits one or more backend lookups, and each lookup has at least four outcomes a test matrix should name — success, no match, timeout, upstream error. The frame below is written with that flow in mind, but every row is a general top-up pattern: no public page tells you how the operator catalogue is stored, who the acquirer is, or how fulfilment is dispatched, and this note stays on the correct side of that line.

Destination and number

Country selection is a small UI with a large downstream effect. The choice sets dial code, number length rules, currency, the operator catalogue the next request will search and often the payment methods on offer. The test rows this generates are mostly data: countries with single national prefixes, countries with mobile-only vs. landline-mobile shared prefixes, countries that share a dial code (US / Canada under +1, Russia / Kazakhstan under +7), territories a visitor will intuitively pick but no catalogue covers, and the empty state when the user has typed nothing yet.

Phone-number validation is where boundary testing earns its name. The format varies by country, numbers can be entered with or without the dial code, visitors paste strings with spaces, dashes, parentheses, non-breaking spaces and leading plus signs copied from a chat app. A matrix covers: minimum and maximum lengths for every country, numbers one digit short, one digit long, non-numeric characters, mobile-prefix mismatches (landline typed when only mobile is supported), E.164 inputs with the dial code typed twice, and the paste-with-spaces case which still trips many flows. For every valid shape there is a UX expectation the test can assert — field accepts, button enables, next step unlocks — and for every invalid shape there is a visible, localised error message that must not be the English fallback.

Operator identification and catalogue

Once the number is in, something behind the scenes decides which operator owns it. The visible outcome falls into four buckets: a confident single operator (the common case), a prefix shared across operators that forces a manual pick, an unknown prefix in a known country, and a timeout from whatever directory service answered. Each should be a row. The catalogue itself is the next axis: operators have different product shapes (fixed amounts, open amounts, bundles, data packs, voice minutes), different currencies on display, and different minimum / maximum ceilings. A good matrix pairs operator with product shape rather than treating them as independent dimensions, because the real failure modes are joins: a prepaid bundle that the operator sells only for numbers already provisioned, or an open-amount top-up blocked because the operator's catalogue flipped to fixed amounts at midnight.

Invalid and boundary inputs

Alongside the data-shape errors above, there is a layer of logical boundaries the visitor can hit: amounts below the operator's minimum, above its maximum, amounts that sit between two bundle tiers, currency conversion rounding at the lowest supported denomination. These are predictable rows. The less-predictable rows are the ones a visitor only hits once: switching country after selecting a product (does the cart clear, keep an orphan line, or try to re-price the old product under the new catalogue?), going back after payment has started, double-submitting the form, and the browser-back case after a success page. Each of these has one correct behaviour and several silent-wrong ones; naming them in a matrix is the only way to prevent regressions.

Checkout and payment outcomes

Payment is the fattest column of the matrix. The visible outcomes are few — paid, declined, abandoned, pending — but they are reached through many branches: card, wallet, local method, 3-D Secure challenge issued, 3-D Secure skipped, soft decline retried on another rail, hard decline, insufficient funds, acquirer timeout, acquirer success with a mismatched reference, chargeback window entry. Every one of these is a test row with an expected final checkout state, a specific error message in the shopper's language, and a specific next action (retry, change method, contact support). Cards also need an expiry-boundary row (the first second of the first day a card is valid, the last second of the last day) and a currency-mismatch row (card issued in one currency, invoice in another). Payment-state tests should assert both the shopper's view and the record the system stores — a success toast with no transaction row is a bug of the loudest kind.

Asynchronous fulfilment and delivery

Fulfilment is the step that separates a top-up from a shopping cart. The money moves instantly; the credit on the destination SIM arrives later, sometimes much later, and arrives through an aggregator we do not control. The matrix needs rows for: immediate success, success within thirty seconds, success within fifteen minutes, success after an hour, success after retry, partial success (operator accepted the request but credited a different amount), and the two failure terminals — permanent failure with a refund and permanent failure pending reconciliation. The user-visible contract on each row is a separate assertion: order status, email, SMS to the sender, and whether the UI promises a refund timeline it can keep.

Retry and idempotency belong here as explicit rows, not as an afterthought. A retried request with the same idempotency key must not debit twice and must not top up twice; a retried request with a reused key for a different amount must be refused; a request that times out at the client but succeeds at the aggregator must be reconciled, not re-sent. The HTTP-state view of this is a general pattern worth asserting independently of the business flow: a 409 on a duplicate submit, a 425 or 409 on an in-flight retry, a 202 on an accepted-but-not-yet-fulfilled transaction, and a stable terminal 200 once the aggregator confirms. Each state has a shopper-facing message the test asserts in every language the site ships.

Confirmation, history and receipts

The post-purchase surface is where trust is retained or lost. The matrix covers the confirmation screen (shows the correct amount, correct destination, correct order id, correct expected delivery window), the email receipt (sent, delivered, and parseable by common clients without broken rendering), the SMS to the sender if the product offers one, and the transaction history page — including filters by status, pagination at the boundaries, and the behaviour of a transaction that moves from pending to failed while the shopper is looking at it.

Localisation, responsive layout, accessibility

Localisation adds a multiplier to every row above: every string the test asserts needs to be asserted in every language the site ships, number formatting follows the destination not the shopper, currency symbols sit on the correct side of the amount, and right-to-left scripts do not break the number input (which usually wants LTR regardless of page direction). Responsive behaviour is simpler but easy to get wrong — the operator picker, the amount grid and the payment form each need rows on narrow viewports, with the on-screen keyboard open and closed, and with the dynamic viewport changing as the browser chrome collapses.

Accessibility is a column of the matrix, not a separate project. Each form field asserts a label association, each error asserts an aria-live announcement, each custom control asserts keyboard operability, and each confirmation surface asserts a focus landing point. For the security layer behind the flow, we anchor coverage to a published reference rather than local opinion: the OWASP Web Security Testing Guide gives test types — input validation, authentication, session management, business logic, client-side — that translate directly into rows in the matrix, so a payment page and a confirmation page can prove they handled the same classes of input the same way.

Running the matrix in CI

A matrix this wide is only useful if it runs on every change. The data-driven approach TestML encourages — one portable test file, one data file, many runtimes — maps well: the countries, operators, products, inputs and expected outcomes live in a table (CSV or YAML), one test file iterates it, and the same file runs under whichever language the service under test is written in. A Python checkout service and a Node fulfilment worker can both prove, from the same source of truth, that a given +44 7700 900000 on operator EE for a £10 product reaches a terminal success state within fifteen minutes. CI pins the matrix at a known revision, flags the row that regressed, and keeps the diff reviewable because the test file does not change — only the data row does.

Writing the matrix once and running it everywhere is the point. The combinations grow faster than any team can click through, but they do not grow faster than a table can describe. Start with the axes above, name the terminal state on every row, and let the runners prove — in every language the stack touches — that a top-up that looked like a single button really did behave like a settlement system.