A/B Testing & Analytics Interview Master Guide
Case Study: Logiqids Quest Feature Activation Project & Statistical Rigor
Target Role: Junior Product Analyst / Domain: EdTech / Core Focus: Experimentation, P-Values & Duration
PM Gamification Mechanics
1. LOGIQIDS PROJECT STORY (STAR FRAMEWORK)
Business Context
Logiqids conducts logical reasoning Olympiads for kids. First stage is free, second stage is paid. Free
worksheets (first 30) act as the activation funnel. However, high drop-off occurred at the top of the funnel
(before completing 10 worksheets).
STAR Phase Executive Breakdown
Situation Observed significant top-of-funnel drop-off. % of newly registered users reaching 10
completed worksheets was low, preventing them from ever reaching the paid worksheet
threshold (Worksheet #30).
Task Drive early activation (% of registered users completing $\ge 10$ worksheets) by introducing
an intrinsic/extrinsic motivation mechanic designed for kids without discounting paid
features.
Action Designed and launched a "Daily Quest" feature (tasks like "Complete 2 worksheets today",
"Maintain a 3-day streak"). Ran a 50/50 A/B test over 30 days on new user sign-ups.
Result Variant group achieved 58% activation rate vs Control's 55% (a 3% absolute lift / ~5.4%
relative lift). Evaluated via Z-test, achieving a p-value $< 0.05$ (statistically significant).
2. ESSENTIAL STATISTICAL TERMS (PLAIN ENGLISH)
Statistical Term Plain English Meaning Interview Pitch Line
Control vs. Variant Control = Old App (55%); Variant = "We ran a 50-50 split test comparing the
App with Quests (58%). Baseline baseline app experience against the Quest
comparison. variant."
Sample Size Total number of users enrolled in the "We calculated our required sample size
experiment. upfront via power analysis to reliably detect
a 2-3% MDE."
P-Value The "Fluke Calculator". Probability "Our p-value dropped below 0.05,
that the 3% lift was pure random luck. confirming less than a 5% chance the lift
was a random fluke."
Statistical Significance The green light! Mathematical proof "Once statistical significance was reached,
that the change caused a real impact. we confidently rolled out Quests to 100% of
users."
Two-Proportion Z-Test The specific math formula used to "Since activation was a binary conversion
compare two conversion percentages. metric (Yes/No), we used a Two-Proportion
Z-Test."
Logiqids A/B Testing & Interview Guide Page 1 of 3
SRM (Sample Ratio Sanity check to ensure the intended "We verified there was no Sample Ratio
Mismatch) 50/50 split was mathematically clean. Mismatch to rule out assignment tracking
bugs."
3. THE 3 FACTORS DICTATING P-VALUE
The p-value is essentially a "signal-to-noise detector". It relies mathematically on three primary levers:
Effect Size (The Gap Size): The larger the gap between Control and Variant (e.g., 55% to 70% vs. 55% to
56%), the easier it is to prove it isn't luck. Larger Lift = Lower P-Value Sample Size (Number of Users):
Larger user volume smooths out random individual quirks. Across 100,000 users, even a 3% gap becomes
undeniable. More Users = Lower P-Value Data Variance / Noise: How unpredictable or noisy the underlying
user behavior is. Lower variance makes the feature signal stand out. Lower Noise = Lower P-Value
4. IMPACT OF A/B TEST DURATION
Duration Type Primary Risks & Trade-offs What Happens to Data Quality?
Too Short • Novelty Effect: Users engage because it's High risk of False Positives. You mistake
($< 14$ days) new, spiking results temporarily. short-term curiosity for a permanent habit
• Day-of-Week Bias: Skews data towards shift.
weekend or weekday habits.
• Funnel Latency: Users don't have calendar
time to reach 10 worksheets.
Sweet Spot • Covers 2–4 full 7-day user cycles. Cleanest Data Quality. Balances statistical
($2 - 4$ weeks) • Novelty effect wears off, revealing true power with operational speed.
habit retention.
• Gives users ample calendar time to
complete funnel milestones.
Too Long • Cookie/Device Churn: Users switch Sample Pollution. Boundaries between
($> 6-8$ weeks) devices/clear cookies, polluting groups. Control and Variant blur over time.
• External Noise: Holidays, marketing
campaigns distort behavior.
• Opportunity Cost: Freezes engineering
velocity while waiting.
5. TOP INTERVIEW CROSS-QUESTIONS & BEST RESPONSES
Q1: Why focus on top-of-funnel (10 worksheets) instead of the paywall at worksheet 30?
"Activation is the leading indicator for monetization. If kids don't build a habit early by solving 10
worksheets, they will never reach the 30th worksheet paywall. Widening the top of the funnel directly
increases the volume of users reaching our monetization gate."
Q2: Why did you run the test for 30 days specifically?
Logiqids A/B Testing & Interview Guide Page 2 of 3
"30 days (about 4 full weeks) hit the optimal balance. It eliminated the initial novelty effect, captured 4 full
weekly cycles of student practice patterns, and allowed enough calendar time for kids to solve 10
worksheets without causing device/cookie pollution."
Q3: What were your guardrail metrics?
"Our main guardrail was Worksheet Completion Accuracy. We monitored score quality to ensure kids
weren't just rushing through or randomly clicking through worksheets just to claim Quest rewards."
6. 30-SECOND INTERVIEW CHEAT SHEET
Executive Pitch:
"In my project at Logiqids, we addressed top-of-funnel drop-off by introducing a gamified Daily
Quest feature. We evaluated it via a 50/50 A/B test over 30 days. Using a Two-Proportion Z-Test,
we measured a 3% absolute lift in activation (55% to 58%) with a p-value $< 0.05$. This
statistical confidence proved the lift was genuine, allowing us to successfully roll out Quests across
our platform."
Logiqids A/B Testing & Interview Guide Page 3 of 3