Files
Magnus HedemarkandGitHub f7d550bb6b fix(site-reliability): make recovery closure gate explicit (#370)
* fix(site-reliability): make recovery closure gate explicit

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): close review gaps

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): link closure evidence sequence

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): link executive closure evidence

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): close authorization and monitoring gaps

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): close mutation and monitoring gaps

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): gate detailed runbook mutations

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): gate remaining operational paths

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): make authorization evidence attributable

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): close final review gaps

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): require independent recovery confirmation

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): close recovery evidence review gaps

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): complete human recovery handoff

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): normalize recovery status tokens

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): require independent resolution approval

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): close authorization consistency gaps

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): normalize incident status guidance

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): carry human confirmation through resolution

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): complete incident closure evidence

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): remove automated recovery claim

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

* fix(site-reliability): normalize monitoring announcement

Signed-off-by: Magnus Hedemark <magnus919@pm.me>

---------

Signed-off-by: Magnus Hedemark <magnus919@pm.me>
2026-08-22 06:27:18 -04:00

32 KiB

Runbook: [Service Name]

Version: [1.0.0] Last Updated: [YYYY-MM-DD] Owner: [Team Name / Individual] Review Cadence: [Quarterly / Bi-annual / Annual]


Table of Contents

  1. Service Overview
  2. SLOs & Error Budget
  3. On-Call Quick Reference
  4. Monitoring & Alerting
  5. Common Failure Modes
  6. Detailed Troubleshooting Procedures
  7. Escalation Paths
  8. Recovery Procedures

Service Overview

Description

[Briefly describe what this service does, its purpose, and the critical function it serves in the broader system architecture.]

Architecture

[High-level description of the service architecture — key components, dependencies, data flow, upstream/downstream services.]

Owner

Field Value
Engineering Team [Team Name]
Team Channel [#slack-channel]
Primary DRI [Name / Role]
Service Catalog URL [https://link-to-service-catalog]
Code Repository [https://github.com/org/repo]
Resource URL
Grafana Dashboard [https://grafana.example.com/d/...]
Datadog Dashboard [https://app.datadoghq.com/dashboard/...]
CloudWatch Dashboard [https://console.aws.amazon.com/cloudwatch/...]
Kibana / Logs [https://kibana.example.com/app/discover#...]
Jaeger / Tempo Traces [https://tracing.example.com/...]
PagerDuty Schedule [https://pagerduty.com/schedules/...]
Status Page [https://status.example.com/]
Runbook (this doc) [link]

Dependencies

Dependency Criticality Notes
[Database e.g. PostgreSQL] Critical [connection details, failover info]
[Cache e.g. Redis] High [cluster info, eviction policy]
[Queue e.g. Kafka] High [topic names, consumer groups]
[External API] Medium [rate limits, quota info]
[Auth provider] Critical [token expiry, rotation schedule]

Key Metrics

Metric Target Description
[p99_latency_ms] < [500ms] End-to-end request latency
[requests_per_second] [N] Throughput
[error_rate] < [0.1%] Ratio of 5xx responses
[cpu_utilization] < [80%] Instance CPU usage
[memory_utilization] < [80%] Instance memory usage

SLOs & Error Budget

Service Level Objectives

SLO Target Window Measurement Method
Latency [99% of requests < 500ms] [28 days] [Histogram buckets / Prometheus]
Availability [99.9%] [28 days] [Ratio of successful requests]
Throughput [Handle N req/s] [1 hour] [Max RPS measured]
Freshness [Data < 5min old] [1 hour] [Lag monitoring]

Error Budget

Period Budget Remaining Burn Rate
28 days [0.1% = 43m 12s] [XX.X% remaining] [Alert if > 2x target]

Burn Rate Alerts

Severity Burn Rate Duration Action
Warning [2x] [1 hour] Investigate
Critical [10x] [6 minutes] Page on-call
Critical [2x] [6 hours] Page on-call

On-Call Quick Reference

The commands below include read-only and mutating examples. Before any command that mutates production, apply the operational closure gate: verify current human authorization for the specific action and scope, record the target, affected population, maximum blast radius, success and abort/rollback criteria, rollback path, and stopping authority. If the gate cannot be satisfied, stop and hand off or escalate.

How to Access

# SSH to production instances
ssh [user]@[bastion-host]
ssh [instance-name].[region].internal

# Kubernetes access
kubectl config use-context [cluster-name]
kubectl get pods -n [namespace]

# Database access
psql -h [host] -U [user] -d [database]

How to Restart

# Restart application service
sudo systemctl restart [service-name]

# Roll restart Kubernetes deployment
kubectl rollout restart deployment/[deployment-name] -n [namespace]
kubectl rollout status deployment/[deployment-name] -n [namespace]

# Safe restart with traffic drain
./scripts/safe-restart.sh [service-name]

Common Commands

# Check service health
curl -s http://localhost:[port]/health | jq .

# View recent logs
journalctl -u [service-name] --since "1 hour ago" -n 100

# Check current version
[service-name] --version
curl -s http://localhost:[port]/version | jq .

# Check active connections
netstat -anp | grep [port] | wc -l

# Check disk space
df -h /data

Common Issues at a Glance

Symptom Try First
Service returning 5xx Check /health endpoint, restart service
High latency on p99 Check CPU/memory, database query times
Alerts firing after deploy Rollback to last known good version
Database connection errors Check connection pool, restart app
Out of memory Increase resources, rollback recent change
TLS / certificate errors Check cert expiry, restart with reload

Monitoring & Alerting

Key Dashboards

  1. Service Overview Dashboard — Primary dashboard for latency, error rate, throughput, saturation (the Four Golden Signals).
  2. Infrastructure Dashboard — CPU, memory, disk, network I/O per instance/container.
  3. Database Dashboard — Connections, query latency, replication lag, cache hit ratio.
  4. Dependency Dashboard — Upstream/downstream health, queue depths, API latency.
  5. Business Metrics Dashboard — User-facing metrics: signups, active users, conversion.

Logging

  • Log Aggregator: [Kibana / Loki / CloudWatch Logs]
  • Log Level: [INFO in production, DEBUG on-demand]
  • Structured Log Format: [JSON]
  • Log Retention: [30 days hot, 90 days cold]
  • Log Query: [{service="[service-name]"} | json]

Useful Log Queries:

# All errors in last hour
{service="[service-name]"} | json | level = "error"

# Requests for a specific user
{service="[service-name]"} | json | user_id = "[user-id]"

# Trace a single request ID
{service="[service-name]"} | json | trace_id = "[trace-id]"

Tracing

  • Tracing Backend: [Jaeger / Tempo / X-Ray]
  • Sampling Rate: [1% head-based, 100% for errors]
  • Trace Query by Service: [service.name="[service-name]"]

Alert Rules

Alert Name Condition Severity Auto-Close
[HighErrorRate] error_rate > [1%] for [5min] Critical [15min after recovery]
[HighLatency] p99_latency > [1s] for [5min] Warning [30min after recovery]
[LowDiskSpace] disk_usage > [90%] Warning [Disabled]
[ServiceDown] up{job="[service]"} == 0 for [1min] Critical [10min after recovery]

Common Failure Modes

Each failure mode is self-contained through mitigation. After any mitigation or recovery action, follow R-01 before declaring the incident resolved.

Before executing any Resolution command, apply the operational closure gate: verify human authorization for the specific action and scope, record the target, affected population, maximum blast radius, success and abort/rollback criteria, rollback path, and stopping authority. These rows describe mitigation options, not permission to execute them. If the gate cannot be satisfied, stop and hand off or escalate.

FM-01: Service Unreachable / High Error Rate

Field Value
Symptom 5xx responses > [N]%, health check failing, pager alert firing
Likely Cause Recent deployment, resource exhaustion, upstream dependency failure
Diagnosis 1. Check /health and /metrics endpoints
2. Review recent deployments (kubectl rollout history or equivalent)
3. Check CPU/memory on instance
4. Check upstream dependencies (database, cache, external APIs)
5. Review recent logs for panic/OOM/panic
Resolution 1. If caused by recent deploy: Rollback immediately: kubectl rollout undo deployment/[deploy]
2. If resource exhaustion: Scale up: kubectl scale deployment/[deploy] --replicas=[N]
3. If upstream failure: Check dependency runbook, pager dependency owner
4. Last resort: Restart the service
5. If none of the above work, escalate
6. After mitigation, follow R-01 before declaring resolved; retain MITIGATING or MONITORING and escalate if evidence is missing.

FM-02: High Latency

Field Value
Symptom p99/p95 latency exceeds SLO threshold, users report slowness
Likely Cause Traffic spike, database query degradation, slow upstream, GC pressure
Diagnosis 1. Check traffic volume vs baseline (is this a spike?)
2. Check database slow query log — SELECT * FROM pg_stat_activity WHERE state = 'active'
3. Check GC metrics (if JVM: jstat -gcutil, if Go: go_memstats_gc_cpu_fraction)
4. Check upstream dependency latencies
5. Review tracing dashboard for slow spans
Resolution 1. Traffic spike: Auto-scale groups should handle; manually increase replicas if needed
2. Slow queries: Kill runaway queries: SELECT pg_terminate_backend(pid) WHERE ...; add missing index
3. GC pressure: Increase heap/memory, tune GC parameters
4. Upstream slow: Circuit-breaker should trip; verify upstream health
5. Temporary fix: Rate-limit or shed non-critical traffic
6. After mitigation, follow R-01 before declaring resolved; retain MITIGATING or MONITORING and escalate if evidence is missing.

FM-03: Out of Memory / OOM Killed

Field Value
Symptom Container/process killed, OOM in kernel logs, instance becomes unresponsive
Likely Cause Memory leak in code, traffic surge, insufficient resource limits
Diagnosis 1. Check `dmesg
Resolution 1. Immediate: Restart the service to reclaim memory
2. Increase limits: Edit resource limits.memory for the container/Pod
3. If caused by code change: rollback the release
4. Schedule memory leak investigation with engineering team
5. Consider enabling memory request-based autoscaling
6. After mitigation, follow R-01 before declaring resolved; retain MITIGATING or MONITORING and escalate if evidence is missing.

FM-04: Database Connection Pool Exhaustion

Field Value
Symptom Application logs show connection refused, too many connections, or connection timeout
Likely Cause Connection leak in application, insufficient pool size, DB restart
Diagnosis 1. Check active connections on DB: SELECT count(*) FROM pg_stat_activity
2. Check max connections: SHOW max_connections
3. Identify connections by application: SELECT application_name, count(*) FROM pg_stat_activity GROUP BY 1
4. Check if connections are idle-in-transaction: SELECT * FROM pg_stat_activity WHERE state = 'idle in transaction'
5. Review application connection pool metrics (hikariCP, etc.)
Resolution 1. Emergency: Kill idle connections: SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE state = 'idle'
2. Kill idle-in-transaction: Same as above with state = 'idle in transaction'
3. Restart the application to reset its connection pool
4. If leak persists, increase max_connections temporarily on DB
5. Schedule fix for connection leak (usually unclosed Connection / Session objects)
6. After mitigation, follow R-01 before declaring resolved; retain MITIGATING or MONITORING and escalate if evidence is missing.

FM-05: Disk Space Full

Field Value
Symptom Disk usage alert firing, application unable to write logs/data, no space left on device
Likely Cause Logs not rotated, data not cleaned up, unexpected large files (core dumps, heap dumps)
Diagnosis 1. df -h to identify full partition
2. du -sh /* 2>/dev/null to find large directories
3. `du -sh /var/log/*
Resolution 1. Immediate: sudo journalctl --vacuum-size=500M or truncate -s0 /var/log/[app].log
2. Clean old logs: find /var/log -name "*.log.*" -mtime +7 -delete
3. Clean temp files: sudo rm -rf /tmp/*
4. Verify logrotate is working: sudo logrotate -f /etc/logrotate.conf
5. If persistent, add disk monitoring alarm at 80%
6. After mitigation, follow R-01 before declaring resolved; retain MITIGATING or MONITORING and escalate if evidence is missing.

FM-06: TLS / Certificate Expiry

Field Value
Symptom Clients receive certificate expired or x509: certificate has expired errors
Likely Cause Auto-renewal failed, cert-manager/ACME issues, manual cert not replaced
Diagnosis 1. Check cert expiry: `echo
Resolution 1. Manual renewal: If cert-manager: kubectl delete certificate [name] to trigger re-issue, or fix the ACME challenge
2. Manual cert replacement: Upload new cert to LB/Ingress
3. Restart ingress controller / LB after cert update
4. If auto-renewal is broken, create a ticket for the platform team
5. After mitigation, follow R-01 before declaring resolved; retain MITIGATING or MONITORING and escalate if evidence is missing.

FM-07: Upstream API Degraded / Down

Field Value
Symptom Our service returns errors for operations that depend on [upstream], upstream latency spikes
Likely Cause Upstream outage, rate limiting, network partition
Diagnosis 1. Check upstream status page
2. Test upstream directly: curl -v https://upstream.example.com/health
3. Check our circuit breaker metrics
4. Check for recent upstream API changes
5. Check network connectivity: ping, traceroute, nslookup
Resolution 1. If circuit breaker open: Wait for half-open / reset, or manually reset if safe
2. If rate limited: Throttle requests, request quota increase
3. If upstream outage: Enable fallback/graceful degradation (serve stale data, queue requests)
4. Page upstream PagerDuty escalation
5. Consider feature flags to disable upstream-dependent features temporarily
6. After mitigation, follow R-01 before declaring resolved; retain MITIGATING or MONITORING and escalate if evidence is missing.

FM-08: Slow Consumer / Queue Backlog

Field Value
Symptom Queue depth growing, consumer lag increasing, messages not processed in time
Likely Cause Consumer crashed or stuck, processing logic regression, queue partition imbalance
Diagnosis 1. Check consumer group lag (Kafka: kafka-consumer-groups --bootstrap-server ... --group [group] --describe)
2. Check consumer process health and logs
3. Check if messages are stuck on poison-pill messages (deserialization errors)
4. Check partition assignment and rebalance events
5. Check downstream that the consumer writes to
Resolution 1. Restart consumers: kubectl rollout restart deployment/[consumer]
2. Skip poison-pill messages: Seek consumer offset past bad message
3. Scale consumers: Increase partitions + consumer replicas
4. If DB is bottleneck: Investigate and resolve DB performance first
5. If backlog is critical, consider replaying messages from an earlier offset after fix
6. After mitigation, follow R-01 before declaring resolved; retain MITIGATING or MONITORING and escalate if evidence is missing.

FM-09: DNS Resolution Failures

Field Value
Symptom lookup [hostname] failures, connection refused, intermittent timeouts
Likely Cause DNS resolver outage, cached stale records, DNS propagation delay, /etc/resolv.conf misconfiguration
Diagnosis 1. Test resolution: dig [hostname] @[dns-server]
2. Check /etc/resolv.conf
3. Check nslookup [hostname] and host [hostname]
4. Check DNS server reachability: nc -zv [dns-server] 53
5. Compare results across different resolvers (e.g., 8.8.8.8)
Resolution 1. Flush DNS cache: sudo systemd-resolve --flush-caches or restart nscd
2. Update resolv.conf: Ensure valid nameservers listed
3. Restart DNS-sidecar (if Kubernetes with CoreDNS/node-local-dns)
4. Override in /etc/hosts as temporary measure
5. If using Kubernetes, check CoreDNS pods and service: kubectl -n kube-system get pods -l k8s-app=kube-dns
6. After mitigation, follow R-01 before declaring resolved; retain MITIGATING or MONITORING and escalate if evidence is missing.

FM-10: Pod CrashLoopBackOff (Kubernetes)

Field Value
Symptom Pod repeatedly restarting, CrashLoopBackOff status, zero ready replicas
Likely Cause Startup failure (config error, missing secret, dependency unavailable), OOM, liveness probe failing
Diagnosis 1. Describe pod: kubectl describe pod [pod-name] -n [namespace]
2. Check pod logs: kubectl logs [pod-name] -n [namespace] --previous
3. Check events: kubectl get events -n [namespace] --sort-by='.lastTimestamp'
4. Verify configmaps/secrets are mounted: kubectl exec -it [pod] -- cat /path/to/config
5. Check liveness/readiness probe configuration
Resolution 1. Fix config/secrets: Correct the ConfigMap or Secret and re-deploy
2. Fix missing dependency: Start the dependency or fix the connection string
3. Increase resources: If OOM-killed, increase limits.memory
4. Fix probe: Correct the probe endpoint/timeout/period
5. Rollback to last known good version if config doesn't help
6. After mitigation, follow R-01 before declaring resolved; retain MITIGATING or MONITORING and escalate if evidence is missing.

Detailed Troubleshooting Procedures

Before executing any command in these procedures that mutates production, apply the operational closure gate: verify current human authorization for the specific action and scope, record the target, affected population, maximum blast radius, success and abort/rollback criteria, rollback path, and stopping authority. These commands are examples, not permission to execute them. If the gate cannot be satisfied, stop and hand off or escalate.

T-01: Initial Incident Triage

                   ┌──────────────────────┐
                   │  Alert Fires / User   │
                   │  Reports Issue        │
                   └──────────┬───────────┘
                              │
                   ┌──────────▼───────────┐
                   │  1. ACKNOWLEDGE      │
                   │  Ack the alert in    │
                   │  PagerDuty/OpsGenie  │
                   └──────────┬───────────┘
                              │
                   ┌──────────▼───────────┐
                   │  2. TRIAGE           │
                   │  - What's affected?  │
                   │  - How many users?   │
                   │  - Is it critical?   │
                   │  - Create incident   │
                   └──────────┬───────────┘
                              │
                   ┌──────────▼───────────┐
                   │  3. INITIAL COMMS    │
                   │  - Post in #incidents│
                   │  - Ping team channel │
                   │  - Update status page│
                   └──────────┬───────────┘
                              │
                   ┌──────────▼───────────┐
                   │  4. DIAGNOSE         │
                   │  Check dashboards,   │
                   │  logs, traces, use   │
                   │  Common Failure Modes│
                   └──────────┬───────────┘
                              │
                   ┌──────────▼───────────┐
                   │  5. MITIGATE         │
                   │  Rollback, restart,  │
                   │  scale up, etc.      │
                   └──────────┬───────────┘
                              │
                   ┌──────────▼───────────┐
                   │  6. VERIFY           │
                   │  Run R-01 closure    │
                   │  evidence checks     │
                   └──────────┬───────────┘
                              │
                   ┌──────────▼───────────┐
                   │  7. RESOLVE          │
                   │  Only after R-01     │
                   │  and stability       │
                   └──────────────────────┘

T-02: Reading Health Check Endpoints

# Standard health check
curl -s http://localhost:[port]/health | jq .

# Verbose health check (includes dependency status)
curl -s http://localhost:[port]/healthz | jq .

# Readiness check (Kubernetes)
curl -s http://localhost:[port]/readyz | jq .

# Expected output for healthy service:
# {"status":"ok","version":"1.2.3","uptime":"72h14m","dependencies":{"database":"ok","cache":"ok","queue":"degraded"}}

T-03: Capturing Diagnostic Data

# Collect a diagnostics bundle
./scripts/diagnostics.sh > /tmp/diag-$(date +%s).txt

# Capture thread dump (Java)
jstack <pid> > /tmp/thread-dump-$(date +%s).txt

# Capture heap histogram (Java)
jmap -histo:live <pid> > /tmp/heap-hist-$(date +%s).txt

# Capture goroutine dump (Go)
curl -s http://localhost:[debug-port]/debug/pprof/goroutine?debug=2 > /tmp/goroutines.txt

# Capture metrics snapshot
curl -s http://localhost:[port]/metrics > /tmp/metrics-$(date +%s).txt

T-04: Safe Rollback Procedure

# Step 1: Identify last known good version
kubectl rollout history deployment/[deployment] -n [namespace]

# Step 2: Rollback to previous revision
kubectl rollout undo deployment/[deployment] -n [namespace]

# Step 3: Monitor rollout status
kubectl rollout status deployment/[deployment] -n [namespace]

# Step 4: Start R-01 closure verification
curl -s http://[service-url]/health | jq .
# A health endpoint is partial evidence. Complete R-01: verify user-facing SLOs
# and critical journeys, dependencies, data/state, secondary effects, and the
# defined stability window and independent human confirmation before declaring the incident resolved.

# Step 5: Keep the incident Mitigating or Monitoring if evidence is missing;
# record the unverified boundary and hand off or escalate.

# Step 6: After R-01 closure evidence and independent human confirmation are
# complete, the human incident owner sends the resolution announcement.

Escalation Paths

Primary Escalation

Level Contact Method Response Time
L1 On-call Engineer PagerDuty / Phone [15min]
L2 [Team Name] Senior Engineer PagerDuty / Slack [30min]
L3 [Team Name] Tech Lead Phone / Slack @mention [1hr]
L4 Engineering Manager Phone / Slack @mention [2hr]

Dependency Escalation

Dependency Team PagerDuty Schedule
[Database Cluster] [Data Platfom Team] [link]
[Kubernetes Cluster] [Infra Team] [link]
[Network / DNS] [Networking Team] [link]
[External API] [Vendor Support] [vendor-ticket-link]

Escalation Procedure

  1. 5 minutes: If not resolved, page L1.
  2. 15 minutes: If L1 cannot resolve, escalate to L2 (Senior Engineer).
  3. 30 minutes: If L2 needs additional context, involve L3 (Tech Lead).
  4. 60 minutes: If incident is customer-facing or SEV-1, notify L4 (Engineering Manager).
  5. SEV-1 criteria: Service down for > 5min, data loss, security breach, revenue impact.
  6. Declare SEV-1: Post in #severe-incidents, create incident channel, invite relevant teams.

Communication Template

INCIDENT: #[incident-number]
STATUS: [Investigating / Mitigating / Resolved]
SERVICE: [service-name]
IMPACT: [Describe user/business impact]
TIMELINE:
  [HH:MM UTC] - Alert fired
  [HH:MM UTC] - On-call acknowledged
  [HH:MM UTC] - Rollback initiated
NEXT STEPS: [What's being done]

Recovery Procedures

R-01: Post-Incident Steps

  1. Verify full recovery at the user boundary — Confirm user-facing SLOs and critical user journeys, relevant dependency health, data/state correctness, and secondary effects such as backlog recovery. Health checks, baseline error/latency, and a cleared alert are partial evidence, not a resolution verdict.
  2. Observe a stability window — Monitor for the defined window and confirm that recovery holds without cascading or delayed effects. Record the evidence and the boundary actually exercised.
  3. Human-confirm the evidence — A human other than the acting automation must review and confirm the complete evidence set, with the confirmation independently attributable to that person. An agent-assigned IC role, automation-authored incident record, green alert, or self-reported health check is not sufficient.
  4. Keep unresolved incidents visible — If any required evidence or human confirmation is missing, retain the incident in MITIGATING or Monitoring, record the unverified boundary, and hand off or escalate rather than marking it resolved.
  5. Update status page — Mark the incident as resolved if used, but only after full recovery evidence and human confirmation are complete.
  6. Resolve alert — Close PagerDuty / OpsGenie alert after the resolution decision, not merely because the alert condition cleared.
  7. Tag, annotate, and notify — Add incident severity, team, and service tags, then send the verified summary to the team channel and affected users.

R-02: Data / State Recovery

Warning: Data recovery procedures should only be attempted by engineers with database admin access.

Before any production restore, verify current human authorization for the specific restore and scope, record the target, affected population, maximum blast radius, success and abort/rollback criteria, rollback path, and stopping authority. Restore to staging and verify integrity first; staging evidence does not authorize the production restore. If the gate cannot be satisfied, stop and hand off or escalate rather than restoring production.

# Step 1: Identify the recovery point (RPO)
# Step 2: Restore from backup to a staging environment first
# Step 3: Verify data integrity in staging
# Step 4: Schedule maintenance window
# Step 5: Perform database restore in production
# Step 6: Verify application works against restored data

Backup Locations:

Resource Backup Frequency Retention Restore Process
[Database] [Hourly WAL + Daily full] [30 days] [link to restore doc]
[Object storage] [Cross-region replication] [N/A] [link to restore doc]
[Configuration] [Git-controlled] [Permanent] Terraform apply

R-03: Incident Report Template

## Post-Mortem: [Title]

**Date:** [YYYY-MM-DD]
**Duration:** [Start] — [End] ([X] minutes)
**Severity:** [SEV-1 / SEV-2 / SEV-3]
**Team:** [Team Name]

### Summary
[2-3 sentence executive summary]

### Timeline
- [HH:MM UTC] — [Event]
- [HH:MM UTC] — [Action taken]
- [HH:MM UTC] — [Recovery verified]

### Root Cause
[What actually caused the incident]

### Impact
- [N] requests failed
- [N] users affected
- [N] minutes of downtime

### Action Items
| Action | Owner | Ticket |
|--------|-------|--------|
| [Fix root cause] | [Name] | [Jira/GH link] |
| [Improve monitoring] | [Name] | [Jira/GH link] |
| [Update runbook] | [Name] | [PR link] |

### Lessons Learned
- What went well:
- What went wrong:
- What we'll do differently:

R-04: Runbook Update Checklist

After any incident, update this runbook:

  • Add any new failure modes discovered
  • Update diagnosis steps that were incorrect or incomplete
  • Update resolution steps that worked
  • Update escalation paths if contacts have changed
  • Update links that were broken
  • Update SLOs / error budget if thresholds have changed
  • Increment version number
  • Update "Last Updated" date

Appendix

A — Useful Scripts & Aliases

# Quick health check
alias svc-health='curl -s http://localhost:[port]/health | jq .'

# Check recent deploys
alias svc-history='kubectl rollout history deployment/[deployment] -n [namespace]'

# Tail logs
alias svc-logs='kubectl logs -f deployment/[deployment] -n [namespace]'

# Shell into a running pod
alias svc-shell='kubectl exec -it deployment/[deployment] -n [namespace] -- /bin/bash'

B — Environment Information

Environment URL / Access Replicas Instance Type
Production [prod-url] [N] [e.g. m5.xlarge]
Staging [staging-url] [N] [e.g. t3.large]
Development [dev-url] [N] [e.g. t3.medium]

C — Configuration Reference

Config Key Description Default Production Value
[DATABASE_URL] Primary DB connection string [redacted]
[REDIS_URL] Redis connection string [redacted]
[LOG_LEVEL] Logging verbosity info info
[MAX_CONNECTIONS] DB pool size 10 [N]
[REQUEST_TIMEOUT] Upstream request timeout 30s 10s
Runbook Service Link
[Database Runbook] [Database] [link]
[Cache Runbook] [Redis/Memcached] [link]
[Infrastructure Runbook] [Kubernetes/AWS] [link]
[Auth Runbook] [Auth Service] [link]

Document Status: This runbook is a living document. If you find errors, omissions, or out-of-date information during an incident, fix it immediately and submit a PR.

Template Attribution: Based on the Site Reliability Engineering principles and Google's SRE books.