Use the IMPACT framework
- Impact: users, journeys, regions and severity.
- Monitoring: detection source and its gaps.
- People: your role, incident commander and collaborating teams.
- Actions: hypotheses, evidence and mitigation sequence.
- Cause: technical and systemic contributors after investigation.
- Transformation: automation, tests, alerts, runbooks and organizational change.
A strong SRE answer
Detection and customer impact
Start with the user-facing symptom, not “CPU was high”. State how you knew and where monitoring was incomplete. If support or customers detected it first, say so and explain the telemetry improvement. Google’s incident guidance recommends timely, actionable alerting centred on user symptoms, with preventive internal alerts for hard limits where appropriate.
Triage and mitigation
Show how you reduced uncertainty: recent change correlation, healthy-versus-unhealthy comparison, dependency signals and a small hypothesis set. Then explain why the mitigation was safe. Rollback, load shedding, feature disablement or failover each has data and dependency risks; naming them demonstrates judgment.
Root cause without blame
Move beyond the triggering line of code. Explain why testing, rollout, monitoring or ownership allowed the issue to reach users. Avoid naming an individual as the cause. A blameless account can still be precise about decisions and control gaps.
MTTR and other metrics
Use timeline metrics you can verify: detection time, declaration, mitigation start, recovery and full resolution. Define what your organization means by MTTR before quoting it. Pair time with impact because a fast recovery can still affect a critical journey.
Communication is operational work
- State who coordinated, operated and communicated.
- Give an update cadence and audience.
- Separate confirmed facts from hypotheses.
- Share impact and workaround before speculative cause.
- Record decisions and timestamps for handoff and postmortem.
- Avoid putting credentials, tokens or sensitive customer records in incident channels or documents.
Preventative automation
Do not end with “we added an alert”. Strong follow-up changes the system: a rollout gate, capacity test, bounded self-healing action, configuration validation, dependency isolation or ownership process. Automation must include preconditions, least privilege, audit, failure handling and rollback. A dangerous one-click production script is not resilience.
Practise more scenarios in 25 Real-World SRE Interview Questions and Troubleshooting Scenarios, How to Transition From System Administrator or DevOps Engineer to SRE.
Conclusion
A credible incident story shows calm prioritization, evidence-based mitigation, clear communication and lasting improvement. The hero is the response system, not one engineer working alone.
Frequently asked questions
Can I discuss a customer incident?
Only within confidentiality obligations. Anonymize the customer and architecture, remove sensitive data and focus on your decisions and learning.
What if I made the mistake?
Own the decision without self-punishment, explain the context and show the concrete change that followed.
Should I give the exact MTTR?
Use it when verified and define the interval. Otherwise describe the timeline without false precision.
What if the root cause was never proven?
Say what evidence established, what remained uncertain and what controls reduced recurrence. Do not turn a hypothesis into fact.
Sources and further reading
Want a second opinion on your profile?
HikeCatalyst can help identify the resume, profile and job-search changes that matter most for your target role.
Get Free Profile Review →