Prompt and setting
An aggregate metric contradicts every important subgroup. The task is to determine whether a segment-mix imbalance, confounder, selection rule, or sampling error created the reversal, and to choose an estimand before applying a statistical fix.
What the interviewer tests
- Recognizing Simpson's paradox as an aggregation reversal, not a mathematical contradiction.
- Checking treatment assignment, segment balance, denominators, and pre-specified analysis plans.
- Separating descriptive association from a causal treatment effect.
Clarifying questions before answering
- Which segments were defined before treatment and which were discovered afterward?
- Was assignment randomized, and are exposure and missing-outcome rates balanced?
- Is the decision about the overall population, a target segment, or a policy effect after standardization?
- Are there sparse cells, repeated users, or post-treatment variables being conditioned on?
30-second answer framework
I would reproduce the aggregate and segment tables with denominators, verify randomization and exposure balance, and identify the variable that changes the segment mix. I would pre-specify the target estimand, then report standardized or stratified effects with uncertainty. If randomization failed, I would rerun or repair the experiment; if the segment is a genuine effect modifier, I would not collapse it into one headline number.
Step-by-step deep dive
1. Rebuild the arithmetic
For every segment, show treatment users, control users, conversions, and rates. Aggregated rates are weighted averages; different weights can reverse the comparison even when every segment favors treatment. Check whether a dashboard used impressions, exposed users, or all assigned users as its denominator.
2. Audit assignment and exposure
Compare treatment probability and segment proportions before outcomes are observed. Check rollout rules, device eligibility, geography, time windows, bots, and missing events. A treatment arm concentrated in a low-baseline-conversion segment can lose in aggregate without contradicting positive within-segment effects.
3. Define the estimand
An overall intent-to-treat effect answers a population launch question. A segment-specific effect answers whether treatment works differently by device. A standardized effect reweights segment-specific estimates to a chosen target population. State which question the decision needs; no adjustment is universally correct.
4. Choose a remedy
With valid randomization and adequate cells, report stratified or regression-adjusted estimates with confidence intervals and test interaction cautiously. With assignment imbalance, fix randomization or rerun. For observational data, draw a causal graph, adjust only for pre-treatment confounders, and use weighting or matching with sensitivity analysis. Never adjust for a post-treatment mediator just to remove the reversal.
5. Communicate the launch decision
Show the aggregate, segment, and standardized estimates together, including uncertainty and segment sizes. Launch globally only when the target estimand and risk tolerance support it; otherwise roll out by segment, collect more data, or pause. Document the analysis plan so repeated subgroup hunting does not inflate false positives.
High-quality sample answer
“I would first rebuild both rates with treatment and control denominators and verify assignment, exposure, eligibility, and missing-event balance. Then I would identify the segment driving the weight difference and define whether we need an intent-to-treat population effect or a segment-specific effect. With sound randomization, I would report stratified and standardized estimates with intervals and inspect interaction; with imbalance, I would rerun or repair the experiment. I would show all views in the launch memo and avoid post-treatment adjustments or unplanned subgroup fishing.”
Common mistakes
- Trust the dashboard aggregate → weighting can reverse the signal → recompute segment numerators, denominators, and weights.
- Declare the segment result causal automatically → assignment or exposure may be biased → audit randomization and missing outcomes first.
- Adjust for every available variable → post-treatment bias can increase → adjust only for justified pre-treatment confounders.
- Pick whichever result supports launch → estimand changes silently → pre-specify the target population and analysis.
Follow-up questions and responses
Is Simpson's paradox always a problem?
No. It signals that aggregate and conditional questions differ. The aggregate may be the right launch estimand, while a segment effect may be the right targeting decision; the issue is failing to state which one is intended.
How do you handle a tiny segment?
Show its uncertainty and avoid overinterpreting a noisy reversal. Consider hierarchical models or a pre-specified pooling rule, and collect more observations before making a segment-specific commitment.
What if the segment variable is post-treatment?
Do not condition on it for the primary causal estimate. Report it descriptively or use a mediation analysis with explicit assumptions; conditioning can create collider or mediator bias.