Mixed-methods prioritization research that replaced opinion with ranked, statistically grounded roadmap inputs across two planning cycles
Survey design · Feature prioritization · Concept testing · Data-quality auditing · Stakeholder alignment
379 + 503
cleaned respondents across two programs
6 → 13
features ranked, framework scaled
p < .01
latent demand detected via pre/post exposure
20.9%
straight-line responders removed before ranking
The problem
Quarterly planning at a top-5 US sportsbook kept running into the same wall: a list of candidate features, several strong opinions about which mattered most, and no shared evidence to settle it. The question I was handed was “which feature should we build?” The question worth answering was “how do we systematically prioritize on user value so this debate doesn’t recur every quarter?”
I built a prioritization framework to answer it, ran it once, and then ran it again a year later at larger scale. The second run is the proof that the first one worked.
Program 1: Personalization and feature prioritization (late 2024)
Design. Twelve unmoderated in-depth discovery interviews across user segments to understand mental models and unmet needs, followed by a prioritization survey of 379 current customers spanning three engagement tiers. Six candidate features were ranked, including a personalization system the product team was most uncertain about.
The method that mattered: pre/post concept exposure. Participants ranked features before and after seeing design concepts. Stated preference alone misses latent demand: people cannot want what they cannot picture. Measuring the shift after exposure separates features that are merely familiar from features that resonate once understood.
Findings.
- Personalization ranked third before exposure and improved significantly after (mean rank 3.66 → 3.41, p < .01). After seeing the concept, 78.1% said they would use it, 63.9% expected it to increase their satisfaction, and 57.3% expected it to increase their share of wallet.
- A feature-awareness gap emerged as the cheapest win in the study: several existing capabilities (live streaming, pre-built parlay editing, search) were valued highly once discovered but largely unknown. Discoverability, not new development, was the fastest path to value.
- A clear hierarchy emerged with cashout and data-visualization improvements first and personalization as a validated follow-on investment.
Outcome. The roadmap sequenced around the evidence rather than the loudest stakeholder, three low-value candidates were not built, and discoverability work moved ahead of net-new development.
Program 2: Promotions and rewards prioritization (late 2025)
Design. A two-wave survey of 636 raw respondents, cleaned to 503 after a data-quality audit removed 20.9% straight-line responders. Thirteen candidate promotion and rewards features were ranked using a rating-plus-ranking approach (more cognitively manageable and more reliable than full ranking of 13 items), with Mann-Whitney U tests between conditions and a composite score weighting multiple measures. The survey was paired with in-person concept testing so the ranking carried qualitative texture, not just numbers.
What changed from Program 1. Scale (13 features vs. 6), the formal data-quality gate, the composite scoring, and the pairing with in-person sessions. The framework itself held: rank on user value, measure before and after exposure, and hand product a defensible order rather than a wishlist.
Findings. The survey produced a clear top tier and a clear bottom tier, and the in-person sessions revealed a comprehension issue with one rewards mechanic that materially changed how its survey score should be read. Reporting that conflict honestly, rather than averaging it away, was the most useful thing the readout did.
Outcome. The ranking became a direct input to the roadmap planning cycle, and the framework was reused a third time for a subsequent parlay-features prioritization.
What I’d tell another researcher
Prioritization research is only as good as its data-quality discipline. Removing a fifth of respondents for straight-lining is uncomfortable to report, but a ranking built on inattentive responses is worse than no ranking at all. The other lesson is about method choice: pre/post exposure is what lets you find the feature people would love but cannot yet imagine, and pairing a survey with in-person sessions is what tells you when a survey score is lying to you.
