Skip to main content
CodingArchitectureintermediate

Production Engineering Runbook & Architecture Documentation

Generate actionable on-call runbooks with telemetry thresholds, mitigation steps, and architectural diagrams.

Compatibility & Specs

Compatible AI Models
ClaudeChatGPTGemini
Last UpdatedOct 3, 2026
Customizable Variables4 parameters

How to Use This Prompt

Follow this 3-step workflow to extract high-signal responses from any compatible AI model.

01

1. Tailor the Parameters

Use the interactive customizer above to substitute the bracketed placeholders with your exact context, requirements, and constraints.

02

2. Send to AI Model

Copy the prompt and paste it into Claude, ChatGPT, Gemini, or Copilot. These models follow structured multi-step constraints reliably.

03

3. Review and Iterate

Review the output against the verified benchmark below. Follow up in the conversation to stress-test edge cases or refine tone.

Prompt Variables & Parameters

Reference breakdown of every dynamic variable embedded in this prompt template.

PlaceholderParameter NameTypeStatusDescription & Guidance
[service_responsibility]Service & ResponsibilitytextRequiredWhat the service is called and what it accomplishesDefault: AuthGuard Service: Go microservice running on Kubernetes handling OAuth2 token validation, API rate-limiting via Redis, and JWT signing for all external API gateways.
[service_dependencies]DependenciestextRequiredUpstream callers and downstream databases/servicesDefault: Upstream: Cloudflare API Gateway; Downstream: Redis Cluster (Rate Limits), PostgreSQL (User Permissions), AWS KMS (JWT Signing Keys).
[critical_alerts]Critical AlertstextareaRequiredPagerDuty alarms that trigger off-hours callsDefault: 1. `AuthGuardHighLatency`: P99 latency > 250ms for 3 consecutive minutes. 2. `RedisClusterFailover`: AuthGuard unable to reach Redis token bucket cache.
[operational_procedures]Operational ProcedurestextareaRequiredKnown recovery commands and maintenance stepsDefault: Kubernetes rolling restart, switching Redis replica to primary, and activating local in-memory fallback rate limiting via LaunchDarkly flag.

Example Execution & Benchmark Output

Sample input arguments and the verified AI response demonstrating expected quality and formatting.

Sample VariablesInputs
[service_responsibility]
AuthGuard Service: Go microservice running on Kubernetes handling OAuth2 token validation, API rate-limiting via Redis, and JWT signing for all external API gateways.
[service_dependencies]
Upstream: Cloudflare API Gateway; Downstream: Redis Cluster (Rate Limits), PostgreSQL (User Permissions), AWS KMS (JWT Signing Keys).
[critical_alerts]
1. `AuthGuardHighLatency`: P99 latency > 250ms for 3 consecutive minutes. 2. `RedisClusterFailover`: AuthGuard unable to reach Redis token bucket cache.
[operational_procedures]
Kubernetes rolling restart, switching Redis replica to primary, and activating local in-memory fallback rate limiting via LaunchDarkly flag.
Expected AI ResponseVerified Result
### On-Call Runbook: AuthGuard Microservice **Service:** `auth-guard` | **Tier:** 1 (Mission Critical) **Team:** Core Security Pod (`#sec-eng-alerts` on Slack) **Dashboard:** `grafana.internal.net/d/authguard-core` ```mermaid graph LR CF[Cloudflare Gateway] --> AG[AuthGuard Go Service] AG --> RC[(Redis Cache - Rate Limits)] AG --> PG[(Postgres - Permissions)] AG --> KMS[AWS KMS - JWT Key] ``` #### Playbook 1: `AuthGuardHighLatency` (P99 > 250ms) - **Business Impact:** External API requests time out; mobile and web logins stall. - **Triage (1st Action):** Check if latency is localized to Redis connection pool: ```bash kubectl logs -n prod -l app=authguard --tail=100 | grep -i 'redis timeout' ``` - **Mitigation (<3 mins):** Enable in-memory local token bucket fallback to bypass Redis network hop: - Navigate to LaunchDarkly -> Flag `authguard-local-ratelimit-fallback` -> Toggle `ON`. - **Verification:** Monitor Grafana p99 panel; latency should drop below 40ms within 60 seconds. #### Playbook 2: Service Restart ```bash kubectl rollout restart deployment/authguard-prod -n prod kubectl rollout status deployment/authguard-prod -n prod --timeout=120s ```

Best Use Cases

Scenarios and roles where this prompt produces maximum leverage.

Engineering teams formalizing on-call runbooks before handing services to rotation
SRE teams standardizing incident playbooks across microservices
Tech leads documenting disaster recovery procedures and failover commands

Tips for Best Results

Techniques to elevate response fidelity

  • •Provide rich background context rather than one-sentence inputs to receive deep, non-generic responses.
  • •Engage in multi-turn conversation: use the initial output as a baseline, then ask the AI to sharpen specific sections.
  • •Prompt the model to highlight any hidden assumptions or missing trade-offs in its recommendations.

Common Mistakes to Avoid

Frequent failure modes and anti-patterns

  • •Giving minimal context and expecting nuanced, expert-level strategic output.
  • •Not validating factual references, citations, or statistical claims with verified primary sources.
  • •Skipping the customization step and pasting raw bracketed template variables into the AI chat.

Part of Curated Collections

This prompt is sequenced as part of these goal-oriented workflows

View all collections

Related AI Prompts

Complementary workflows in Coding

View all Coding prompts
Codingintermediate

Legacy Codebase Refactoring & Technical Debt Reducer

Systematically deconstruct tightly coupled spaghetti code into isolated, testable pure functions without breaking behavior.

claudechatgptgemini
#refactoring#technical-debt#code-smells
Codingadvanced

Architectural Decision Record (ADR) Analysis & Trade-Off Matrix

Evaluate competing technical architectures, state assumptions, and document a formal Architectural Decision Record.

claudechatgptgemini
#system-architecture#adr#trade-offs
Codingadvanced

Production Bug Forensic Root-Cause Analysis & 5-Whys

Conduct a blameless post-mortem, trace crash telemetry, and execute a 5-Whys root cause investigation.

claudechatgptgemini
#root-cause-analysis#post-mortem#debugging

Related Engineering Guides

Deep-dive playbooks and system prompt methodologies for Coding

View all guides