Field notes on staying on-call.
Guides on escalation, rotations, and incident response — for engineering teams, MSPs, healthcare, and families.
Escalation Policies: How Many Steps, and How Long to Wait Between Them
Most escalation policies fail for the same two reasons: too few fallback responders, and wait times tuned for a calm afternoon instead of 3am. Here's how to fix both.
Read the postWriting a Runbook Your On-Call Engineer Will Actually Use at 3am
A runbook that reads well in a design review often fails completely at 3am. Here's what changes when you write for someone who's half-asleep and under pressure.
Read the postWhy "Repeating Until Acknowledged" Beats a Single Push Notification
A phone on silent will happily eat a single push notification without a trace. Repeating, escalating alerts are what actually close that gap.
Read the postOn-Call Burnout: Early Warning Signs and How Rotation Design Prevents It
Burnout doesn't start with someone quitting — it starts with paging patterns that are visible months earlier, if you know where to look.
Read the postOn-Call Alerting for Healthcare Teams: What to Look for Beyond Compliance Theater
A HIPAA badge on a vendor's homepage tells you almost nothing about whether their alert will actually wake someone up. Here's what to check instead.
Read the postAlert Triage for MSPs: Handling Multiple Clients Without Burning Out Your Team
One client's "critical" is another client's "check it in the morning." MSPs need a triage layer that most alerting tools weren't built for.
Read the postUsing a Paging App for Elderly Care and Family Medical Emergencies
A missed-call notification and a paging alert are not the same thing. When it's a fall detector or a medical alert button, that difference matters.
Read the postHow Small Teams Handle On-Call Compensation
There's no universal standard for paying on-call — but there are a handful of models small teams keep converging on, for good reasons.
Read the postWebhook-Based Alerting: Connecting Prometheus and Grafana Without an Integration Project
Most monitoring tools already speak webhook. The "integration" is usually just pointing them at the right URL.
Read the postWhy Teams Are Leaving Per-User Pricing for Incident Alerting
Per-user pricing looks fine at 5 people. The math changes fast once a team is adding its 8th or 12th engineer to the rotation.
Read the postA Practical Incident Postmortem Template for Small Teams
The postmortem template that gets used consistently is the one that takes fifteen minutes, not the one with forty sections.
Read the postWhat Is Alert Fatigue, and How Do You Actually Reduce It?
Alert fatigue isn't a mindset problem — it's a direct, measurable result of your signal-to-noise ratio, and it's fixable.
Read the postOn-Call Schedule Best Practices for Small Teams
A rotation that works for a 20-person SRE org will burn out a 4-person team fast. Here's what actually fits a small team.
Read the post