A two-test experimental study across three prototypes that identified the only layout to perform on both tasks while structurally preventing accidental spends of rewards currency
Experimental design · Latin Square counterbalancing · Unmoderated testing · Error prevention · Responsible gaming
62
participants, 126 prototype exposures
0% vs 4.8%
accidental-purchase rate, safest vs. riskiest layout
68.8%
first-choice preference for the winning layout
+1.16
composite desirability (−3 to +3)
The problem
In the betslip, users can apply a free profit boost or one purchased with rewards currency. If a design leads someone to spend rewards currency when they meant to use a free boost, that is an accidental purchase the user did not intend, and a direct harm to trust. The team had three candidate layouts and needed to know which best helped users find and apply the correct boost while preventing that error.
The design: rigor that shipped in two weeks
Two complementary tests, run simultaneously on an unmoderated platform:
- Interaction test (quantitative), n=32, within-subjects with Latin Square counterbalancing. Every participant saw all three prototypes; the Latin Square controlled for order effects given the platform’s randomization limits. Captured task success, first-click accuracy, errors, and comprehension (did the user know which boost type they had selected?).
- Think-aloud test (qualitative), n=30, between-subjects, 10 per prototype. Clean first impressions of a single prototype each, with comparison screenshots shown only at the end to gather cross-prototype preference without contaminating initial reactions.
- Non-directive task wording by design. Task 1 said only “apply a 25% profit boost,” never “free.” That tests whether the design itself communicates the free-vs-paid distinction, so accidental rewards-currency selections surface as the design flaw they are.
- Also piloted a standardized desirability scale for reuse across future work.
What I found
Scroll position dominated behavior. Whichever boost appeared first was selected at near-perfect rates. Prototype C (free first) hit 97.6% on the free task but only 66.7% on the paid task; Prototype B (paid first) hit 86.4% on the paid task but 57.1% on free. Each single-scroll layout optimized one task at the other’s expense.
Only the separated layout (Prototype A) performed well on both (66.7% free / 77.5% paid), the balanced choice.
Error severity was the real story. In Prototype B, 4.8% of users were charged rewards currency without realizing it: an accidental purchase. In Prototypes A and C, that harmful-error rate was 0%. Prototype A’s separation gave users clarity even when they erred: those who selected the paid boost knew it was the paid option.
Preference matched safety. 68.8% ranked Prototype A first in the interaction test and 53.3% chose it in the think-aloud test, citing the clear separation. “I don’t want to accidentally spend my rewards” captured the sentiment.
A separate applied-state problem surfaced. The checkmark was ambiguous (confirm vs. already applied), causing users to re-tap. I recommended a distinct applied indicator and a clear removal control.
The recommendation and why
Ship Prototype A’s separated layout. It is the only design balanced across both tasks, it structurally prevents accidental rewards-currency purchases, and it is the design users prefer. When the safest option is also the most preferred, the trade-off conversation gets easy. The redesign scored +1.16 composite on the piloted desirability scale (−3 to +3).
What I’d tell another researcher
This is where methodological care pays off directly. A single-prototype test, or directive task wording, would have hidden the accidental-purchase risk entirely; it only appears when you let people act on an ambiguous design and then check whether they understood what they did. Designing the study to expose the harmful error, not just measure task success, is what made the recommendation trustworthy. It also ties to responsible gaming: preventing unintended spending is a user-protection outcome, not just a usability one.
