Question and Context
An on-demand delivery marketplace wants to test a new dispatch algorithm across 24 operating zones. A treatment order can claim a courier that a control order would otherwise receive, and couriers can move between adjacent zones. Design an experiment to estimate the effect of rolling the algorithm out to an entire market. Explain the estimand, randomization unit, spillover and carryover controls, power calculation, analysis, guardrails, and launch decision.
The 24 zones, intervention, and operating details are interview-case assumptions, not facts about a real platform. Assume the algorithm is reversible and can be selected by an experiment service. Do not assume a switchback window or washout period; those must be justified from order and courier dynamics.
This is a hard experiment-design question for data scientists, experimentation scientists, economists, and product analysts. A public interview-question page classifies a network experiment as a hard data-science and data-analyst case and records an asking date of April 7, 2026. LinkedIn has documented cluster experiments for detecting social-network interference. DoorDash has documented region-time switchbacks for dispatch and, in a 2025 marketplace-ads case, explained why carryover and low power can make switchbacks a poor choice when a resource can instead be isolated. An ICLR 2026 paper continues to study the bias-variance problem in cluster-randomized network experiments. Together, these sources support current representativeness and technical relevance. They do not establish how frequently any company asks the question or justify assigning it to a named company.
The central failure is to randomize orders and compare two user-level means. The algorithm changes access to a shared courier pool, so one order's treatment can change another order's outcome. That violates the no-interference condition behind an ordinary A/B interpretation. A strong answer defines the effect of interest first, maps how treatment travels through space and time, and then makes the assignment, estimator, uncertainty calculation, and rollout decision answer that same question.
What the Interviewer Is Evaluating
First, can the candidate define the causal question? “Does treatment beat control?” is incomplete. The product decision concerns the effect of applying the new algorithm to an entire market, which can differ from the direct effect on one treated order while neighbors remain mixed.
Second, can the candidate describe interference as a mechanism? Shared couriers create competition between orders; courier movement connects adjacent zones; open deliveries and behavioral responses create carryover across time. Naming “network effects” without identifying these edges does not determine a design.
Third, can the candidate choose among designs rather than reciting “cluster randomization”? Geographic clusters reduce cross-arm competition but can leave few independent units. Switchbacks create more region-time units but can be biased by carryover. Two-stage saturation designs can separate direct and spillover effects in a social graph. Resource splitting can create isolated universes when the scarce budget or inventory can actually be partitioned.
Fourth, can the candidate align analysis with assignment? Millions of orders do not create millions of independent observations when treatment is assigned to a few market-time blocks. Power, weighting, standard errors, and randomization checks must reflect the units and assignment probabilities in the design.
Finally, can the candidate reject a misleading result? A strong answer checks assignment, exposure, boundary crossings, persistence, concurrent experiments, and metric maturity before reading lift. It also states when a switchback estimates a short-run policy effect but cannot support a claim about long-run equilibrium or learned courier behavior.
Clarifying Questions to Ask
- Which effect drives the decision? Is the target the effect per eligible order if every zone uses the new policy, the average effect per market-time block, a direct effect on treated orders, or spillover onto untreated orders? The estimator and weighting depend on this choice.
- Where does interference travel? Can one courier serve multiple zones? Can an order be matched across a boundary? Do pricing, incentives, batching, or queue priorities also couple the zones? The answer defines the interference graph.
- How long can treatment affect later outcomes? What are the distributions of assignment, pickup, completion, courier repositioning, and return-to-supply times? A reversible code path does not imply an immediately reversible marketplace state.
- How many independent markets exist? Are the 24 zones separate labor pools or subdivisions of four tightly connected cities? Twenty-four labels may represent only four useful geographic clusters.
- Can the scarce resource be isolated? A budget, candidate inventory, or quota may be split into independent pools. Couriers usually cannot be duplicated, so resource-universe randomization may be infeasible for dispatch.
- What is the primary outcome and denominator? Completed deliveries per eligible order, pickup time, courier utilization, and gross bookings answer different questions. Eligibility must be fixed before treatment can change queue entry or completion.
- Which delayed and two-sided guardrails matter? Candidate guardrails include cancellations, late deliveries, courier earnings per active hour, idle time, acceptance, customer complaints, and cross-zone displacement.
- What concurrent policies change the same market? Promotions, pricing, weather responses, and other dispatch experiments can interact with assignment. They may require blocking, exclusion, or an explicitly factorial design.
- What assignment and exposure logs exist? The team needs the scheduled variant, actual algorithm, market cluster, time block, order eligibility time, courier transitions, fallback, and outcome maturity to distinguish design failure from product effect.
30-Second Answer Framework
“I would first define the estimand as the change per eligible order if the whole target market used the new dispatch policy. Order-level randomization is invalid because treatment and control compete for the same couriers. I would build a movement and matching graph, combine zones with heavy cross-boundary traffic, and, because the policy is reversible, evaluate a randomized cluster-by-time switchback. The block and any fixed transition window would come from observed carryover, not a default duration. I would plan power from historical cluster-block variance and autocorrelation, analyze intention to treat using the assignment schedule, and use an estimator and uncertainty method that match the per-order estimand and clustered design. Before reading lift, I would verify assignment, actual exposure, contamination, and metric maturity. I would ship through a persistent regional ramp only if completion improves and customer and courier guardrails pass; long-lived learning or unresolved spillover would require a longer cluster rollout instead.”
Step-by-Step Deep Dive
Step 1: Define the estimand before choosing the bucket
Write the target contrast in words: “For orders that would be eligible under either policy, what is the change in the primary outcome when every unit in the target market follows the new algorithm rather than the current algorithm?” This is a full-rollout policy effect.
Under interference, an order's outcome depends on its own assignment and the assignments around it. A mixed-market user-level test primarily estimates a direct effect under partial saturation; it need not equal the all-treatment versus all-control contrast. If the business wants both direct and spillover effects, name them as separate estimands rather than expecting one treatment-control difference to identify both.
Freeze eligibility before the policy can affect it. For example, define eligible orders from request validity, service area, and a timestamp recorded before dispatch. An outcome such as completed deliveries per eligible order preserves failures and unmatched orders. An “average delivery time among completed treatment orders” conditions on a post-treatment event and can make a policy that drops hard orders look faster.
Step 2: Turn “network effects” into an interference map
Create edges from operational data and product rules:
- orders are connected when they compete for the same courier during overlapping intervals;
- zones are connected by cross-zone matching and courier movement;
- time blocks are connected while an assigned delivery remains open or a courier remains displaced;
- policies are connected when pricing, batching, incentives, or another experiment changes the same capacity.
Use pre-experiment data for this map. Do not build clusters from behavior created by the treatment. Quantify the share of matches and courier transitions crossing each proposed boundary. A useful clustering keeps most strong edges within a cluster while retaining enough clusters for power. LinkedIn's published approach makes the same trade-off explicit: more clusters provide more randomization units, while larger clusters reduce edges crossing between treatment and control.
The map also identifies outcomes the design cannot isolate. If couriers regularly cross every zone boundary, zone-level assignment only relabels a connected market. The team may need to randomize at the whole-city level over time, use separate persistent markets, or accept a weaker causal claim.
Step 3: Compare designs against the mechanism
Individual Bernoulli randomization has high nominal sample size and is useful when units do not affect one another. It is inappropriate for this dispatch policy because treatment orders consume shared capacity and contaminate control.
Persistent geographic cluster randomization assigns all eligible traffic in a sufficiently isolated market cluster to one policy for the full experiment. It handles long carryover and behavioral learning better than rapid switching, but 24 zones may collapse into only a few connected markets. Balance, power, and generalization then become difficult.
Clustered switchback randomization assigns a connected geographic cluster to one policy for a time block, then rerandomizes future blocks. It fits a reversible dispatch algorithm whose important effects decay within a measurable period. DoorDash's published dispatch design uses region-time units for this reason. It creates more experimental units, but neighboring blocks are still connected when open orders, relocated couriers, or changed behavior persist.
Two-stage or saturation randomization first assigns a network cluster to a treatment saturation, then assigns members within it. It can identify how outcomes change with neighbors' exposure and is useful for social or referral features. It is more complex and less natural when one dispatch policy must govern the same shared courier pool at a moment.
Isolated-universe randomization partitions the interfering resource itself. DoorDash's 2025 ads example splits campaign budgets so variants cannot consume one another's budget. This can be stronger than a switchback when the resource is divisible and each universe can operate normally. It does not create a valid solution by pretending the same courier belongs to two independent pools.
For the stated case, choose a clustered switchback only after confirming that the main effect decays quickly enough. Group heavily connected adjacent zones into market clusters, then randomize market-cluster × time-block units. If the effect includes weeks of courier learning or persistent repositioning, use a persistent cluster rollout instead.
Step 4: Design the switchback schedule and transition rule
Estimate carryover from pre-experiment order lifetimes, courier transitions, and operational simulation, then test candidate schedules with A/A data. The block should be long enough that a large share of the relevant state resolves, but shorter blocks create more units and usually more power. Region size creates the same trade-off: larger regions reduce spatial leakage and reduce the number of independent units.
Randomize every block within prespecified strata such as market cluster, day of week, and time-of-day band. Do not deterministically alternate treatment and control, because a fixed sequence can align with demand trends. Balance sequences and prevent a cluster from receiving only one variant during a critical time band.
Define transitions before outcomes arrive. The primary analysis can attribute every eligible order to the policy scheduled at its eligibility time and follow it to maturity. If the protocol uses a boundary washout, make that interval fixed by the schedule, exclude it from both arms, explain the changed estimand, and include its loss in power. Do not remove “contaminated-looking” orders after inspecting outcomes.
A compact predeclared assignment record could be:
experimental unit = market cluster × time block
assignment = randomized within market cluster × time-of-week stratum
eligibility time = recorded before dispatch policy selection
primary outcome = completed delivery within the fixed maturity horizon per eligible order
primary estimand = full-rollout effect per eligible order
transition rule = fixed before exposure; no outcome-based exclusionsStep 5: Plan power from effective units, not order count
Use historical A/A periods or replayed assignment schedules to estimate variance across cluster-blocks, serial correlation within a market, common time shocks, unequal traffic, and any planned covariate adjustment. Re-randomize the proposed schedule many times and measure how often it detects a prespecified minimum effect under realistic carryover assumptions.
The effective sample is driven by independent assignment information, not by the raw number of deliveries nested inside a block. Shortening blocks can add units but increase carryover bias. Enlarging clusters can reduce spatial leakage but remove units. Power planning should show this bias-variance frontier and select the smallest effect worth the operational cost, not search for a schedule that guarantees significance.
Prespecify duration, allocation, stopping rules, and any variance reduction using only pre-treatment covariates. If available clusters cannot identify the minimum decision-worthy effect, state that limitation before launch. A giant order count cannot repair four independent markets.
Step 6: Analyze according to the design and estimand
Use intention-to-treat assignment from the randomized schedule. Report compliance and actual algorithm exposure separately; do not delete fallback orders from treatment. Aggregate diagnostics at cluster-block level, because that is where assignment changes.
Match weighting to the estimand. An unweighted mean across cluster-blocks estimates an average block effect. If the decision target is an effect per eligible order under rollout and cluster sizes vary, use a design-consistent Horvitz-Thompson or Hájek-style estimator based on known assignment probabilities, and state its target population. Report the average block effect as a secondary view rather than silently switching estimands.
For uncertainty, prefer randomization inference over the actual assignment scheme when feasible. A model-based alternative must account for repeated observations within markets, serial dependence, shared time shocks, and stratification. Treating every order as independent produces artificially narrow intervals. Report the estimate, interval, minimum decision-worthy effect, and sensitivity to the prespecified transition rule.
Before impact, run health checks:
- scheduled treatment shares and balance within each stratum;
- assignment persistence and actual algorithm exposure;
- eligible-order, fallback, and missing-event rates by arm;
- cross-cluster matching and courier movement;
- outcome maturity and delayed cancellations;
- concurrent experiment and policy overlap;
- A/A calibration of false-positive rates under the same pipeline.
Step 7: Tie the result to a rollout decision
Choose one primary outcome tied to the dispatch objective, such as completed deliveries within a fixed horizon per eligible order. Add customer guardrails such as cancellation, lateness, complaint rate, and pickup wait; courier guardrails can include earnings per active hour, idle time, acceptance, and unsafe workload concentration. Define harm margins and rollback thresholds before exposure.
A positive switchback result supports a short-run market policy effect under the tested blocks. It may not capture equilibrium supply, courier learning, or retention over weeks. Follow it with a persistent regional ramp or holdout when those dynamics matter. Compare the direction and magnitude with the switchback, monitor boundaries, and stop expansion if guardrails or contamination exceed the predeclared limit.
If cluster and individual experiments disagree, do not average them into one number. The difference can be evidence of interference, as LinkedIn's published meta-experiment illustrates. Revisit each estimand, cluster leakage, assignment health, and power before deciding which design better matches full rollout.
High-Quality Sample Answer
“I would start by making the causal question explicit. The launch decision is the change in completed deliveries per eligible order if the entire target market uses the new dispatch algorithm. That is different from treating one order while neighboring orders remain on control. I would define eligibility before dispatch so the treatment cannot improve its own denominator, and I would follow every assigned order to a fixed maturity horizon.
Order-level randomization is not valid here. A treated order can take the only suitable courier and change the control order's wait or completion, while courier movement spreads that effect across zone boundaries and later time periods. I would use pre-experiment matching and movement data to construct an interference graph, then combine adjacent zones until cross-cluster matching and courier movement fall below a predeclared diagnostic threshold. If the 24 zones reduce to only a few connected markets, I would say so rather than claim 24 independent clusters.
Because the dispatch code is reversible, my first candidate is a clustered switchback. Each connected market cluster receives one algorithm for a randomized time block. I would randomize blocks within market and time-of-week strata, not alternate deterministically. Window length comes from order completion and courier repositioning data plus A/A schedule simulations. Short windows provide more units but raise carryover; long windows reduce leakage but lose power. If the policy changes courier behavior for weeks, I would reject the switchback and use a persistent market-level rollout.
The primary analysis is intention to treat using the scheduled algorithm at the order's pre-dispatch eligibility time. I would keep fallback and failed assignments in their original arm. Because assignment occurs at market-block level, power and uncertainty come from those units and their serial dependence, not millions of orders. I would simulate the proposed schedule on historical A/A data, prespecify the minimum effect, duration, and transition rule, and use design-based randomization inference where feasible. Since the decision metric is per eligible order, I would use a design-consistent weighted estimator with known assignment probabilities and also report the unweighted average market-block effect.
Before reading lift, I would check treatment balance within strata, actual policy exposure, fallback and missing-event rates, cross-cluster matches, courier transitions, outcome maturity, and overlapping experiments. My primary outcome would be completed deliveries within the fixed horizon per eligible order. Customer cancellations, lateness, complaints, pickup wait, and courier earnings, idle time, and acceptance would have predeclared harm margins.
If the estimate clears the minimum useful effect and all guardrails, I would run a persistent regional ramp to verify that the short-run switchback result survives learning and equilibrium changes. If carryover remains material, clusters are too few for the planned effect, or individual and cluster estimates disagree without a clear mechanism, I would not present the user-level result as a rollout effect. I would redesign the experiment or narrow the causal claim.”
Common Mistakes
- Randomizing millions of orders independently → Treatment and control consume the same courier supply, so the nominal sample size hides interference → Randomize a unit that contains the main interaction or isolate the resource itself.
- Saying “use clusters” without defining the edge → Geography may not contain courier movement, social interaction, budget competition, or temporal carryover → Build the interference map from the actual mechanism and pre-treatment data.
- Choosing a standard 30-minute switchback → A convenient window may be shorter than delivery and repositioning effects or longer than needed → Estimate carryover and calibrate candidate schedules with A/A data and simulation.
- Alternating variants deterministically → Treatment can align with rush periods, weekdays, or trends → Rerandomize blocks within prespecified time and market strata.
- Counting every order as independent → Assignment happens at a coarser unit, producing understated uncertainty and false confidence → Base power and inference on the actual randomized structure.
- Reporting an unweighted block mean as a per-order rollout effect → Unequal block sizes change the target population → Define the estimand and use a design-consistent estimator with explicit weighting.
- Dropping boundary orders after seeing results → Outcome-driven exclusions break the randomized comparison → Predeclare eligibility, transition handling, and maturity, then keep intention-to-treat assignment.
- Treating switchback as universally robust → Carryover, learning, low power, and common time shocks can invalidate or weaken it → Use persistent clusters, saturation, or isolated universes when their assumptions fit better.
- Shipping from a short-run lift alone → Marketplace equilibrium and courier behavior may change under persistent rollout → Validate with a staged persistent ramp and two-sided guardrails.
Follow-Up Questions and Responses
Follow-up 1: What if the 24 zones collapse into only four independent markets?
Four market clusters provide weak information for a persistent cluster experiment, and many orders inside them do not solve that. If effects decay quickly, repeated randomized market-time blocks may add information, provided serial dependence and carryover are modeled. If effects last for weeks, the honest options are a longer paired or phased market rollout, stronger pre-treatment covariate adjustment, simulation for operational risk, or waiting for more independent markets. Report wide uncertainty and narrow the claim; do not manufacture independence by splitting connected zones.
Follow-up 2: What if treatment changes courier behavior for several weeks?
A rapid switchback estimates a mixture of current policy and history and is a poor design for the long-run effect. Use persistent cluster assignment or a randomized phased rollout, measure activation, hours supplied, earnings, retention, and customer outcomes over the full behavioral horizon, and keep assignment at the market level. A switchback can still screen immediate algorithmic effects, but it cannot replace the long-horizon experiment.
Follow-up 3: How would the design change for a social sharing feature?
Build a graph from pre-treatment relationships or interaction edges. Cluster randomization can approximate an all-treated versus all-control network contrast when most edges remain inside clusters. If the goal includes direct and spillover effects, randomize clusters to different treatment saturation levels and then randomize users within clusters. Prespecify exposure mappings such as the treated share of relevant neighbors, and plan for lower power and overlap between communities.
Follow-up 4: What if the interference channel is a divisible advertising budget?
Split the budget or inventory into independent universes and give each variant exclusive access to its pool. This removes direct cannibalization while retaining user-level assignment and can improve power. Validate that pools cannot borrow, pacing and auctions do not reconnect them, and the split represents the full-rollout policy. If isolation is incomplete, measure leakage and keep the causal claim bounded.
Follow-up 5: What if individual-level and cluster-level experiments disagree?
First verify assignment, exposure, metric definitions, maturity, and power in both. Then compare the estimands: an individual test measures treatment in a mixed environment, while a cluster test moves more neighbors together and may better approximate rollout. Quantify cross-cluster edges and saturation, and use a prespecified meta-experiment or randomization test when possible. A meaningful difference is evidence to investigate interference, not a reason to select the larger lift.
Follow-up 6: How would you choose the switchback window in practice?
Plot the decay of open orders, courier displacement, and key metrics after policy transitions using pre-experiment data or safe pilots. Simulate several randomized schedules on A/A periods, measuring false-positive calibration, variance, and sensitivity to fixed boundary buffers. Select the shortest window whose residual carryover stays below the predeclared tolerance while power remains adequate, then run sensitivity analyses on neighboring window and transition choices.