Designing an On-Call Rotation That Works
The engineer who resolves incidents fastest ends up on-call more often, because managers route the hard pages to whoever will fix them quickest — and the rotation quietly stops being a rotation for that person. This is the single most common failure mode in on-call design, and it is self-reinforcing: the more reliably someone resolves incidents, the more incidents get routed their way, the less bandwidth they have for the work that got them promoted in the first place, and eventually they leave. A rotation that punishes competence is not a rotation; it is an informal tax on your best responders.
Why the obvious fix doesn’t work
The obvious fix — “just enforce the rotation strictly, no exceptions” — runs into a real problem: some incidents genuinely need the person who understands that subsystem, and refusing to page them “because it’s not their week” produces a slower, worse resolution and a customer-facing outage that lasts longer than it needed to. The rotation and the reality of uneven system knowledge are in real tension, and pretending otherwise just moves the tax underground: the on-call engineer pages the expert informally, off the books, and the expert absorbs the interruption anyway with none of the compensation or acknowledgment a formal page would carry.
What actually works: separate the rotation from the escalation path
The fix is to stop treating “who is on-call” and “who gets paged for the hard cases” as the same question. A working design has three explicit tiers:
- Primary on-call — a strict, enforced rotation. This person owns triage, the first response, and anything within documented runbooks. No exceptions for “but they know it better,” because the primary’s job is triage and communication, not necessarily deep fixes.
- Secondary on-call — backup for the primary, paged automatically if the primary doesn’t acknowledge within a set window. Same strict rotation.
- Named escalation contacts per subsystem — explicitly not a rotation. These are the people with deep knowledge of a specific area, paged only when the primary’s triage determines the incident needs them, and every such page is logged and counted toward that person’s load explicitly, not absorbed silently.
The critical design decision is tier three being visible and counted. If the expert gets pulled in three times this month, that is data that should show up in a load-balancing conversation, a comp conversation, or a “we need a second person who understands this subsystem” conversation — not disappear into “well, they’re just really good at this.”
A comparison of rotation designs
| Design | Who bears the load | Failure mode |
|---|---|---|
| Strict single-tier rotation, no escalation path | Whoever is on-call this week, evenly | Hard incidents take longer to resolve; primary is stuck without the right context |
| No formal rotation, “page whoever knows it” | The most competent engineers, disproportionately | Burnout concentrated on your best people; invisible to management until they quit |
| Primary/secondary rotation + logged named escalation | Primary/secondary evenly; specialists only when genuinely needed, and visibly | Requires more process discipline to log escalations consistently |
A worked example
A payments team had one engineer, call her the person who wrote the reconciliation service, who ended up on every reconciliation incident regardless of whose week it was — informally, over Slack, because paging the actual on-call engineer “would just mean they page her anyway.” Six months in, she was effectively on-call every week while carrying a full project load, and nobody outside the team could see it because none of those pages went through the formal system. The fix was mechanical: reconciliation incidents became a named escalation path, triaged by the strict primary rotation first, and every escalation to her was logged as a paged event with the same tooling as a primary page. Within a quarter, the data showed the actual load, which became the business case for training a second engineer on that subsystem — something that had been “obviously a good idea” for a year with no evidence to force the prioritization.
A checklist for auditing your current rotation
- Pull six months of pages and check whether any single person’s informal interruption count (Slack pings, “hey can you look at this”) dwarfs their formal on-call load. If so, you have a hidden tax.
- Confirm your rotation tooling actually enforces the strict tiers — if a manager can still directly message a specific engineer instead of paging through the system, the rotation is advisory, not real.
- Check whether named escalation contacts exist for your highest-incident subsystems, and whether pages to them are logged the same way as primary pages.
- Ask your most reliable responders directly whether they feel like they’re “off the hook” during their off weeks. If the honest answer is no, the design is broken regardless of what the schedule says.
- Use the escalation-frequency data explicitly in headcount and training conversations — it is a direct, defensible signal for where a team has a single point of failure.
Limitations
This model adds process overhead: someone has to actually log escalations, and a team culture that resists “bureaucracy” will resist logging a two-minute Slack answer as a formal page. It also does not fix an underlying problem where a system genuinely has only one person who understands it — the rotation design surfaces that problem clearly, but surfacing it is not the same as solving it; solving it requires investing in cross-training, which is a separate and slower commitment.
FAQ
Doesn’t logging every escalation just add friction to getting quick help? A little, and that’s an acceptable cost. The alternative — an untracked tax on your best engineers — is worse and, because it’s invisible, never gets fixed until someone burns out or leaves.
What if the team is too small for three tiers? Collapse secondary and escalation into one tier, but keep primary separate and strictly rotated, and still log any off-rotation help as an event. The principle (measure the informal load) matters more than the exact tier count.
How does this relate to measuring engineering impact more broadly? It’s a specific case of a general problem covered in Measuring Engineering Impact: work that doesn’t show up in the metrics you’re tracking gets systematically undervalued until it’s measured directly.
Bottom line
An on-call rotation that lets the same competent engineers get informally paged outside their rotation is not actually distributing the load — it’s taxing your best people invisibly until they burn out. Separate the strict primary/secondary rotation from named, explicitly logged subsystem escalation, and use the resulting data to make the case for cross-training before you lose the person who was quietly carrying the system. See also Hiring Engineers Who Ship and Remote-First Engineering Teams for related team-design tradeoffs.
Related Articles
Remote-First Engineering Teams That Deliver
Dipankar Sarkar explains how to build remote-first engineering teams that outperform co-located ones: communication architecture, async workflows, and trust.
Hiring Engineers Who Ship: A Practical Guide
Dipankar Sarkar shares the hiring framework used to build teams at Nykaa, Hike, and Orangewood Labs: identifying engineers who ship, not just interview well.
Observability for AI Agents in Production
Standard APM tells you a request was slow. It doesn't tell you why an agent picked one tool over another. Here is what to log, trace, and alert on instead.