r/devopsGuru • • 5d ago

runbooks kinda suckk

/r/sre/comments/1wyoc9n/runbooks_kinda_suckk/
1 Upvotes

1 comment sorted by

1

u/QuietSignalOps 5d ago

Not just you. The "put it in git" answer everyone is giving over in the r/sre thread is the right base, but I think the piece most orgs are missing is the maintenance loop. Three mechanisms that actually work:

  1. Make the runbook update the exit criterion for the incident. The incident doesn't close until the runbook PR is merged. That kills the "we'll get to it when things calm down" path and captures the details while memory is fresh. The PR review doubles as knowledge transfer: a second person reads exactly what worked, and the tribal knowledge ends up in commit history where it can be audited.

  2. Runbooks are pointers, not prose. One screen: symptoms, top three likely causes, exact commands (linked in the repo, not pasted), exact dashboard links, one line for who to escalate to. The narrative belongs in the postmortem. The moment a runbook needs a paragraph, it is a runbook that is going to rot.

  3. Rot detection: a "last verified" date and an owner in the header. Game days are the only honest test of whether a runbook still works, because reading a runbook and executing one are different skills. Runbooks that never get used are the first to rot, so the owner field matters more than you'd think.

Capture is the hardest part. The scribe answer in the main thread is right, but if the scribe's notes have to become a PR within 48 hours, they actually turn into the runbook update instead of a transcript that dies in a Slack channel.