The Customer Communication Playbook for Incidents
Customer communication during incidents is the discipline that separates companies that survive outages from the ones that lose accounts because of them. The technical fix matters. The communication around it matters just as much. Customers who are informed promptly, honestly, and with a clear timeline tolerate downtime far better than customers left in silence.
What you actually need to know
- The first update should go out within 15 minutes of confirming an incident, even if you have no answers yet.
- Customers tolerate downtime. They do not tolerate silence during downtime.
- Status page plus email is the minimum communication stack. Notifications inside the app add significant reach.
- The post incident review is a document meant to build trust, not a defensive one. Write it to rebuild, not to protect.
- Every communication during an incident should include: what is affected, what you are doing, and when the next update will come.
| Communication Channel | Reach | Latency | Best For |
|---|---|---|---|
| Status page | Users who check | Immediate | All incidents, first channel |
| Email to affected accounts | High | 5 to 10 minutes | Major incidents, account level impact |
| In app banner or notification | Active users only | Immediate | Products with high active session rates |
The core argument
The worst incident communication I have seen is the one that does not exist. An outage that lasted 45 minutes where the company posted nothing until after service was restored. By the time they updated the status page, the support inbox had 200 tickets, three enterprise customers had emailed their account managers, and one of them had started a cancellation request.
The outage itself was recoverable. The silence was not. Customers who are left in the dark during an outage fill the silence with assumptions. Usually the worst ones. A company that does not communicate during a major incident is either hiding something or does not have its act together. Neither impression is good.
The playbook is simple. Acknowledge fast. Update regularly. Explain honestly. Commit to a follow up. The specific words matter less than the cadence. A status update every 30 minutes during a major incident tells customers that someone is working the problem and that they will not be forgotten.
The incident communication template
Initial acknowledgment (within 15 minutes): "We are currently investigating an issue affecting [service or feature]. We know this is impacting [number or type] of customers. Our team is on it. Next update in 30 minutes."
Progress update (every 30 minutes during the incident): "Update as of [time]: We have identified [what you know]. We are working on [what you are doing]. Service is expected to recover by [time estimate, or 'we will update in 30 minutes']."
Resolution announcement: "As of [time], the issue affecting [service] has been resolved. [Brief description of what was wrong and how it was fixed]. We will publish a full post incident review by [time]."
Post incident review (within 24 hours): Timeline, root cause, mitigation steps, prevention actions. One to two pages. Published on the status page and sent to affected customers.
Common mistakes teams make during incidents
- Waiting to communicate until you have all the answers. The first update does not need answers. It needs acknowledgment.
- Posting one update and then going silent for an hour. The cadence matters as much as the content.
- Being vague about impact. "Some users may experience issues" is less useful than "Users attempting to log in from 9:15 AM to 9:45 AM UTC were unable to complete authentication."
- Writing the post incident review as a legal document. Customers want a plain English explanation, not indemnification language.
- Not publishing a post incident review at all for major incidents. Customers who experienced significant downtime are watching to see if you learned from it.
Where to start: a three step incident communication setup
Step 1: Set up your status page before you need it. Use Statuspage.io, BetterUptime, or a self hosted option. Write the templates for initial acknowledgment, progress updates, and resolution announcements ahead of time. Do not write templates under pressure.
Step 2: Define the communication roles. Who updates the status page during an incident? Who sends the email? Who manages incoming support tickets? Define these roles before the incident. Confusion about who is responsible for customer communication during an active incident wastes critical time.
Step 3: Run a tabletop drill. Once per quarter, simulate an incident response. A team member plays the customer. The on call team works through the incident communication protocol. The first real incident should not also be the first time you have practiced the communication workflow.
FAQ
Frequently asked
- How quickly should I send the first incident update?
- Should I publish updates on the status page or send email?
- How honest should I be about the cause during the incident?
- What goes in the post incident review that customers see?
- How do I write the post incident summary without making my team look bad?
Author
Closing note from the author
I keep these closing notes short on purpose. Most engineers writing about this topic are not the engineer you want to hire. I might be. Yashveer Singh, founder of Yashveer Labs. The contact channel is Instagram. The proof is the portfolio. The standard is in the work. If we are aligned, you will know within five minutes of the first message.