Skip to main content
CodingDebuggingadvanced

Production Bug Forensic Root-Cause Analysis & 5-Whys

Conduct a blameless post-mortem, trace crash telemetry, and execute a 5-Whys root cause investigation.

Compatibility & Specs

Compatible AI Models
ClaudeChatGPTGemini
Last UpdatedOct 3, 2026
Customizable Variables3 parameters

How to Use This Prompt

Follow this 3-step workflow to extract high-signal responses from any compatible AI model.

01

1. Tailor the Parameters

Use the interactive customizer above to substitute the bracketed placeholders with your exact context, requirements, and constraints.

02

2. Send to AI Model

Copy the prompt and paste it into Claude, ChatGPT, Gemini, or Copilot. These models follow structured multi-step constraints reliably.

03

3. Review and Iterate

Review the output against the verified benchmark below. Follow up in the conversation to stress-test edge cases or refine tone.

Prompt Variables & Parameters

Reference breakdown of every dynamic variable embedded in this prompt template.

PlaceholderParameter NameTypeStatusDescription & Guidance
[incident_timeline]Incident Timeline & ImpacttextareaRequiredWhen it happened, what broke, customer impactDefault: 10:14 AM: Marketing sent email campaign to 500k users. 10:18 AM: Web dashboard latency spiked from 150ms to 24,000ms. 10:22 AM: PostgreSQL database CPU hit 100% and connection pool exhausted. 10:40 AM: Database restarted, traffic throttled.
[error_telemetry]Stack Trace & TelemetrytextareaRequiredError messages, Datadog alerts, or SQL slow query logsDefault: Postgres log: `LOG: duration: 8412.314 ms statement: SELECT * FROM user_notifications WHERE user_id = $1 ORDER BY created_at DESC;` Node.js error: `TimeoutError: ResourceRequest timed out after 10000ms at Pool.acquire`
[system_code]System Code / QuerytextareaRequiredThe code or query involved in the incidentDefault: The notification dropdown component in the web navigation bar triggers `GET /api/notifications` on every page navigation, executing `SELECT * FROM user_notifications WHERE user_id = $1 ORDER BY created_at DESC LIMIT 10`.

Example Execution & Benchmark Output

Sample input arguments and the verified AI response demonstrating expected quality and formatting.

Sample VariablesInputs
[incident_timeline]
10:14 AM: Marketing sent email campaign to 500k users. 10:18 AM: Web dashboard latency spiked from 150ms to 24,000ms. 10:22 AM: PostgreSQL database CPU hit 100% and connection pool exhausted. 10:40 AM: Database restarted, traffic throttled.
[error_telemetry]
Postgres log: `LOG: duration: 8412.314 ms statement: SELECT * FROM user_notifications WHERE user_id = $1 ORDER BY created_at DESC;` Node.js error: `TimeoutError: ResourceRequest timed out after 10000ms at Pool.acquire`
[system_code]
The notification dropdown component in the web navigation bar triggers `GET /api/notifications` on every page navigation, executing `SELECT * FROM user_notifications WHERE user_id = $1 ORDER BY created_at DESC LIMIT 10`.
Expected AI ResponseVerified Result
### Forensic Root-Cause Analysis (RCA) #### 1. The 5-Whys Investigation 1. **Why did the API return 504 Timeouts?** Node.js backend exhausted its database connection pool waiting for queries to complete. 2. **Why was the connection pool blocked?** Each connection was tied up for 8.4 seconds executing `SELECT * FROM user_notifications`. 3. **Why did that query take 8.4 seconds?** The database performed a Sequential Table Scan across 42 million notification rows. 4. **Why was there a sequential scan?** The index on `user_notifications` only indexed `(user_id)`, forcing an in-memory Sort of hundreds of notifications per user before applying `LIMIT 10`. 5. **Why was this index gap not caught in staging?** Staging had only 1,000 rows where in-memory sorting executed in 1ms, masking the O(N log N) scaling cliff. #### 2. The Permanent Fix 1. **Database:** Add compound index `CREATE INDEX CONCURRENTLY idx_user_notifications_composite ON user_notifications (user_id, created_at DESC);` 2. **Application:** Cache user unread notification counts in Redis with a 60-second TTL instead of querying Postgres on every client page load.

Best Use Cases

Scenarios and roles where this prompt produces maximum leverage.

SREs and tech leads conducting blameless post-mortem reviews after production outages
Developers diagnosing complex performance degradations under heavy traffic spikes
Engineering managers preparing transparent root-cause reports for executive stakeholders

Tips for Best Results

Techniques to elevate response fidelity

  • •Provide rich background context rather than one-sentence inputs to receive deep, non-generic responses.
  • •Engage in multi-turn conversation: use the initial output as a baseline, then ask the AI to sharpen specific sections.
  • •Prompt the model to highlight any hidden assumptions or missing trade-offs in its recommendations.

Common Mistakes to Avoid

Frequent failure modes and anti-patterns

  • •Giving minimal context and expecting nuanced, expert-level strategic output.
  • •Not validating factual references, citations, or statistical claims with verified primary sources.
  • •Skipping the customization step and pasting raw bracketed template variables into the AI chat.

Part of Curated Collections

This prompt is sequenced as part of these goal-oriented workflows

View all collections

Related AI Prompts

Complementary workflows in Coding

View all Coding prompts
Codingadvanced

Pragmatic REST & GraphQL API Contract Architect

Design robust, backwards-compatible API contracts with clean error schemas, pagination, and idempotency keys.

claudechatgptgemini
#api-design#rest-api#graphql
Codingadvanced

Relational & NoSQL Database Schema & Indexing Audit

Audit database tables, composite indexes, foreign keys, partition strategies, and N+1 query vulnerability points.

claudechatgptgemini
#database-schema#sql-optimization#indexing
Codingadvanced

End-to-End & Integration Test Boundary Specification

Architect integration test boundaries with ephemeral test containers, database seed strategies, and external API stubs.

claudechatgptgemini
#integration-testing#e2e#testcontainers

Related Engineering Guides

Deep-dive playbooks and system prompt methodologies for Coding

View all guides