Skip to main content
CodingDebuggingadvancedFeatured

Production Incident Post-Mortem & Root-Cause Synthesizer

Convert messy incident Slack logs and alerts into a blameless, rigorous post-mortem with corrective action items.

Compatibility & Specs

Compatible AI Models
ClaudeChatGPTGemini
Last UpdatedApr 1, 2026
Customizable Variables3 parameters

How to Use This Prompt

Follow this 3-step workflow to extract high-signal responses from any compatible AI model.

01

1. Tailor the Parameters

Use the interactive customizer above to substitute the bracketed placeholders with your exact context, requirements, and constraints.

02

2. Send to AI Model

Copy the prompt and paste it into Claude, ChatGPT, Gemini, or Copilot. These models follow structured multi-step constraints reliably.

03

3. Review and Iterate

Review the output against the verified benchmark below. Follow up in the conversation to stress-test edge cases or refine tone.

Prompt Variables & Parameters

Reference breakdown of every dynamic variable embedded in this prompt template.

PlaceholderParameter NameTypeStatusDescription & Guidance
[incident_severity]Incident Severity & Affected ServicetextRequiredSeverity level and impacted serviceDefault: SEV-1: Customer Payment Processing Outage
[downtime_duration]Downtime / Degradation DurationtextRequiredTotal elapsed time from first error to full resolutionDefault: 34 minutes full checkout outage; 14 minutes elevated error rates
[raw_incident_notes]Raw Incident Notes & Slack LogstextareaRequiredPaste unformatted incident notes, alerts, and chat snippetsDefault: 14:12 UTC - Automated Datadog alert: Stripe webhook consumer backlog exceeds 10,000. 14:16 UTC - On-call engineer notices Redis connection pool saturated (100% capacity). 14:22 UTC - Deploy pipeline locked by an unrelated PR merge. 14:29 UTC - Redis maxclients reached because worker pods auto-scaled from 6 to 48 without connection limits. 14:38 UTC - Hard restart of worker pods with lowered pool sizes restores processing. 14:46 UTC - Backlog cleared, error rate returns to 0.01%.

Example Execution & Benchmark Output

Sample input arguments and the verified AI response demonstrating expected quality and formatting.

Sample VariablesInputs
[incident_severity]
SEV-1: Customer Payment Processing Outage
[downtime_duration]
34 minutes full checkout outage; 14 minutes elevated error rates
[raw_incident_notes]
14:12 UTC - Automated Datadog alert: Stripe webhook consumer backlog exceeds 10,000. 14:16 UTC - On-call engineer notices Redis connection pool saturated (100% capacity). 14:22 UTC - Deploy pipeline locked by an unrelated PR merge. 14:29 UTC - Redis maxclients reached because worker pods auto-scaled from 6 to 48 without connection limits. 14:38 UTC - Hard restart of worker pods with lowered pool sizes restores processing. 14:46 UTC - Backlog cleared, error rate returns to 0.01%.
Expected AI ResponseVerified Result
### 1. Executive Summary On October 12, 2026, from 14:12 UTC to 14:46 UTC (34 minutes full outage, 14 minutes degraded recovery), our payment processing service experienced a complete inability to ingest customer checkout webhooks. Approximately 1,420 transactions were delayed, with zero corrupted financial data. The root cause was an un-capped database connection pooling configuration that exhausted Redis maxclients when background worker pods auto-scaled eightfold during a traffic surge. ### 3. Root Cause Analysis (The 5 Whys) 1. **Why did checkout fail?** The payment worker pods threw timeout exceptions trying to acquire Redis lock connections. 2. **Why were connections unavailable?** Redis reached its hard ceiling of 10,000 concurrent client sockets. 3. **Why were there 10,000 sockets?** Horizontal Pod Autoscaler scaled workers from 6 to 48 pods, and each pod initialized a static pool of 200 persistent connections (48 x 200 = 9,600 + system connections). 4. **Why did each pod initialize 200 connections?** Connection pool sizing was hard-coded in 2024 when the cluster ran only 4 fixed pods. 5. **Why was this not caught in staging?** Staging traffic never reaches thresholds required to trigger pod autoscaling beyond the minimum replica count (Systemic flaw: lack of load-test scale parity). ### 5. Corrective Action Items - **[P0]** Implement Redis client connection pooling limits tied dynamically to pod replica counts (Owner: SRE | Due: Oct 15). - **[P1]** Add Datadog threshold alert when Redis client count exceeds 70% of maxclients (Owner: Infra | Due: Oct 16).

Best Use Cases

Scenarios and roles where this prompt produces maximum leverage.

Engineering teams drafting formal post-mortems after production service outages
SREs standardizing incident management and corrective action tracking across squads
Tech leads presenting outage post-mortems to non-technical executive stakeholders

Tips for Best Results

Techniques to elevate response fidelity

  • •Provide rich background context rather than one-sentence inputs to receive deep, non-generic responses.
  • •Engage in multi-turn conversation: use the initial output as a baseline, then ask the AI to sharpen specific sections.
  • •Prompt the model to highlight any hidden assumptions or missing trade-offs in its recommendations.

Common Mistakes to Avoid

Frequent failure modes and anti-patterns

  • •Giving minimal context and expecting nuanced, expert-level strategic output.
  • •Not validating factual references, citations, or statistical claims with verified primary sources.
  • •Skipping the customization step and pasting raw bracketed template variables into the AI chat.

Part of Curated Collections

This prompt is sequenced as part of these goal-oriented workflows

View all collections

Related AI Prompts

Complementary workflows in Coding

View all Coding prompts
Codingadvanced

Principal Code Reviewer & Architecture Auditor

Conduct rigorous architectural code reviews identifying memory leaks, race conditions, and typing holes.

claudechatgptcopilot
#code-review#clean-code#architecture
Job Interviewadvanced

Behavioral Failure & Workplace Conflict Answer Architect

Structure answers for difficult behavioral questions about interpersonal conflict, failed projects, and bad decisions.

claudechatgptgemini
#behavioral-questions#conflict-resolution#failure-story
Codingintermediate

Exhaustive Edge-Case Unit & Integration Test Generator

Analyze production functions to discover subtle concurrency, boundary, and null pointer edge cases and write unit tests.

claudechatgptgemini
#unit-testing#integration-testing#edge-cases

Related Engineering Guides

Deep-dive playbooks and system prompt methodologies for Coding

View all guides