Behavioral Interview: How Do You Communicate During an Incident with Incomplete Facts?
Prompt and scope
Tell me about a production incident where facts were incomplete but you had to update customers, support, or leadership. How did you decide what to say now, what to withhold, and how to revise your judgment when new evidence appeared?
This question tests judgment under pressure, ownership boundaries, and communication habits rather than memorization of a status-page template. GitLab’s incident guidance calls for regular updates that describe customer impact and mitigation, while coordinating with incident responders and the incident lead before public communication. A strong answer shows how you make uncertainty legible without allowing an information vacuum to amplify anxiety.
What the interviewer evaluates
- Separating confirmed facts, hypotheses, and unknowns.
- Confirming customer impact before choosing audience, channel, and cadence.
- Giving an action, owner, and next-update time instead of only reporting a problem.
- Correcting the record when evidence changes, without hiding the early judgment.
- Protecting sensitive information and avoiding personal blame for an unverified cause.
- Demonstrating that communication reduced repeated questions, wrong mitigations, or trust loss.
Clarifying questions to ask
- Was the impact internal, limited to one customer, or a public-service incident?
- Were you the incident lead, technical responder, or communication coordinator?
- What action did each audience need to take?
- Which facts were verified by monitoring, logs, or responders, and which were hypotheses?
- Did security, privacy, or compliance constraints limit disclosure?
30-second answer
I first confirm the impact and the facts I am authorized to represent. I split each update into knowns, unknowns, current actions, and the next update time. External messages contain verifiable customer impact and mitigation, never a guessed cause or personal blame. Technical responders continue investigation while I maintain a severity-appropriate cadence and tailor the action field for support and leadership. If new evidence changes the story, I publish a clear correction, explain the impact, and review whether the communication helped people take the right next step.
Step-by-step deep dive
1. Confirm role and customer impact
Identify the incident lead, the approver for public updates, and the responder who can provide technical facts. Start with which customers are affected and how; if impact is still being verified, state the check and its owner.
2. Classify the information
Label information as confirmed fact, working hypothesis, unknown, or next verification time. “Some requests returned 5xx for 20 minutes” is a fact; “the database pool is exhausted” is a hypothesis until verified. This tells listeners which parts may change.
3. Write actionable versions for each audience
Customers need impact, a safe workaround, and the next update. Support needs recognition cues, approved wording, and escalation paths. Leadership needs scope, business risk, resource requests, and decision points. Include technical detail only when it enables action, and keep logs and personal data out of public channels.
4. Set cadence and one source of truth
Keep a timeline or shared incident document with message versions, evidence, owners, and send times. Set a severity-based cadence; if nothing material changed, say that investigation continues, no new impact is confirmed, and when the next update will arrive. The communication lead coordinates cadence while responders validate content.
5. Handle uncertainty and new evidence
Use qualifiers such as “currently confirmed,” “being verified,” and “not observed,” rather than turning an unknown into a negative claim. If evidence changes scope or mitigation, correct the message promptly and state what changed, why, and whether customers must act.
6. Resolve disagreement and protect sensitive data
If engineers want to wait for root cause while support needs an immediate notice, propose the smallest verifiable update and let the incident lead approve it. Route security, privacy, and customer-specific details through restricted channels; keep public text to necessary impact and action.
7. Prove value through outcomes and review
Record whether updates were on time, repeated support questions fell, customers took the right mitigation, and any promise was wrong. A blameless review should inspect information sources, approval paths, and templates for delay, then assign measurable improvements.
High-quality sample answer
During a payment-callback delay, I coordinated support and engineering responders. We first confirmed that some merchants exceeded the callback SLA, but not the cause. I split the timeline into confirmed impact, a queue hypothesis under investigation, current mitigation, and a 30-minute update time. Customers received the delay range, a temporary assurance against duplicate charges, and the next notification. Support also received recognition cues and an escalation path.
Engineering later confirmed that a consumer deployment caused the backlog. I had the incident lead review the evidence, then corrected the public scope and identified the newly affected merchants. After recovery, I reviewed on-time updates, duplicate tickets, and customer missteps, and added consumer-lag alerts and a named public-update approver. Customers knew when to expect information, and the team did not mislead them with a guessed cause.
Common mistakes
- Guessing the cause to sound certain → later corrections erode trust → label facts, hypotheses, and unknowns.
- Saying only “we are investigating” → people lack impact and next steps → provide impact, action, owner, and time.
- Waiting for full root cause → customers and support invent their own story → publish the smallest verifiable impact.
- Copying internal technical detail to customers → confusion or disclosure risk → rewrite for audience and action.
- Silently editing an earlier update → people may rely on the old version → publish a dated correction with impact.
- Blaming the messenger in review → people hide uncertainty → inspect systems, process, and approval paths.
Follow-up questions and responses
Would you message before customer impact is confirmed?
First verify impact quickly. If an internal audience must wait, state the check, owner, and next update time. A public message needs enough evidence to support its claims; vague wording should not manufacture certainty.
What if engineers want root cause but support wants an immediate notice?
Turn the disagreement into a minimal verified update: observable impact, mitigation in progress, and the next update time. Have the incident lead approve it; root cause can follow later.
When must you correct an early message?
Immediately when scope, customer action, recovery state, or risk changes materially. State the change, its reason, and whether customers must do anything again.
How do you avoid conflicting versions from different teams?
Maintain one timeline and message owner. Support, leadership, and public channels use the same evidence and adapt only the action fields for the audience.
What if the incident involves security or privacy?
Bring in security, legal, and privacy owners and restrict the channel according to data classification. Public text should contain only approved impact and action, not investigation details.
How do you prove communication was effective?
Measure on-time updates, repeated questions, wrong mitigations, support tickets, and post-recovery trust feedback. Convert failures into process changes with an owner and deadline.