Latest / Tech Leadership with Fexingo: Engineering Managers, CTOs, and Technical Leadership Conversations / Why Your On-Call Rotation Is Destroying Team Morale
Transcript
- Lucas: If you ask most engineers what they hate most about their job, on-call duty is right up there with meetings that could have been an email. Luna: I think the worst part isn't the 2AM pages, it's the creeping dread that your weekend might get wrecked by a system you didn't even touch. Lucas: Exactly. And the standard fix — rotating the burden evenly across the team — actually makes things worse. I want to talk about why, and what a better model looks like. Luna: Are you about to tell me the dedicated on-call team approach? I've seen that work at a few places, but it's controversial. Lucas: Yeah, I am. And controversial is the right word. But let me start with a concrete case. A mid-sized SaaS company I worked with — about 80 engineers, running a B2B platform with strict uptime SLAs — had the standard weekly rotation. Every six weeks you were primary on-call for seven days. And every six weeks, that engineer produced about half their normal output. They were constantly context-switching, interrupted during deep work, and dreading the rotation. Luna: So the burnout was baked into the system. They thought fairness meant everyone shares the pain, but really it just spread the damage. Lucas: Right. And the worst part was that the most junior engineers were the most anxious, but they were also the least equipped to handle a complex outage. So the senior engineers ended up being shadow-on-call anyway, because juniors would escalate everything. The ostensible fairness was a mirage. Luna: What did they do instead? Lucas: They created a dedicated incident response team. Four engineers, rotating on a three-month stint, fully focused on reliability. No feature work. They owned the pager, the runbooks, and the postmortems. Everyone else was off the hook for on-call completely. Luna: I can already hear the pushback — 'that's a demotion, it's boring, you're taking away their chance to build things.' Lucas: Sure, that's the common objection. But the counterintuitive truth is that a focused reliability tour can be a huge growth opportunity. You learn the system end to end, you develop crisis management skills, and you become the go-to expert. Plus, when you rotate back to feature work, you bring that deep systems knowledge. Luna: Right, and the company I saw do this had a rule: the dedicated team was always three people, never fewer. That way you have coverage for PTO and sickness without forcing overtime. Lucas: That's smart. The company I mentioned started with four, and it worked well. They also made sure the stint was exactly three months — long enough to get good at it, short enough to avoid burnout from the intensity. Luna: What were the results? Lucas: After six months, they measured a 40 percent drop in self-reported burnout scores from the engineers who had been through the rotation. Incident response time — mean time to acknowledge — dropped by 30 percent, because the response team was constantly monitoring and could react faster. And the feature teams actually shipped more: uninterrupted deep work blocks went up by nearly 25 percent. Luna: So the math is clear: the cost of having four engineers not shipping features was more than offset by the productivity gain from everyone else. Lucas: Exactly. And the other hidden benefit: the incident response team wrote better runbooks, built better dashboards, and automated away the most common pages. After a year, the number of alerts per week dropped by half, because they were fixing the systemic causes instead of just firefighting. Luna: That's the virtuous cycle. The dedicated team makes the system more reliable, which reduces the load on the next team, who can then improve it further. Lucas: Now, I should say — this approach isn't right for every team. If you have fewer than, say, 15 engineers, a dedicated team might be too expensive. But there are smaller-scale adaptations. You can have a rotating 'incident commander' role that's separate from the primary on-call, so one person coordinates while someone else actually debugs. Luna: Or you can schedule on-call in shorter blocks. I've seen some teams do 24-hour shifts instead of a full week, so the dread period is shorter. Lucas: That helps, but it doesn't solve the context-switching problem. The real insight is that on-call is fundamentally interrupt-driven work, and interrupt-driven work is the enemy of deep work. If you treat on-call as a tax that everyone pays equally, you're just distributing the cost of a broken system rather than fixing it. Luna: So the goal should be to minimise the total tax, not to distribute it fairly. Lucas: Exactly. And that's what the dedicated team model does. It concentrates the tax on a small group for a limited time, and uses the leverage of that focused attention to reduce the overall tax for everyone. Luna: I think a lot of CTOs resist this because it feels like admitting that your system is unreliable. But every system is unreliable at some level. Lucas: True. And the alternative — pretending that everyone can be equally productive while carrying the pager — is just denial. I'd rather see an engineering leader say, 'Our system needs active management, and we're going to resource that properly.' Luna: What about the career development angle? I've heard engineers worry that time on the dedicated team will hurt their promotion prospects. Lucas: That's a real concern, and it's up to management to set the right incentives. The company I mentioned included reliability work in their promotion criteria explicitly. They had a separate track for 'reliability engineer' that was equivalent to senior feature engineer. And they made sure that graduating from the stint came with a visible credential — a presentation to the whole engineering org, a written retrospective that went into the performance review. Luna: So it was treated as a leadership rotation, not a punishment. Lucas: Exactly. And the engineers who went through it often came out as better architects, because they understood failure modes deeply. That's something you can't get from building features. Luna: One thing I want to push back on: the dedicated team model can create a knowledge silo. If only four people understand the incident response process, what happens when they rotate out? Lucas: Good point. The solution is rigorous documentation and running postmortems that are shared broadly. The company I mentioned had a weekly 'incident review' meeting open to all engineers, where the response team walked through the week's notable events. And the runbooks were living documents, updated after every significant incident. By the end of a three-month tour, the team had usually improved the runbook significantly, making it easier for the next team to take over. Luna: So you're building institutional knowledge while also reducing the overall burden. Lucas: Right. And the shared review meeting also helped spread the learning without everyone having to live through the incident. It's like the difference between being a firefighter and studying fire science. Both are valuable, but you don't want every engineer to have to fight fires to learn how fires work. Luna: Let me play devil's advocate one more time. What about the cost? Four engineers fully dedicated to reliability is a big line item. How do you justify that to a CFO? Lucas: You frame it in terms of opportunity cost. What's the cost of 80 engineers operating at 50 percent productivity for one week out of six? That's about 13 full time equivalent engineers worth of lost output per year. So the dedicated team of four is actually a net gain of nine FTEs. Plus, faster incident response reduces customer churn. The CFO math works if you do it right. Luna: Yeah, when you put it that way, it's a no-brainer. But most engineering leaders don't calculate the hidden cost of the rotation. Lucas: They don't. And that's partly because the cost is invisible — it's lost deep work, increased context-switching, and the anxiety tax. Those things are hard to measure but they're real. Luna: I think a good first step for a team that's not ready for a full dedicated team is to measure the true cost. Track how many hours per week each engineer loses to on-call interruptions, whether they're paged or not. Lucas: That's a great idea. Even just asking engineers to log their 'on-call overhead' for a month can be eye-opening. I've seen teams discover that the person who's on-call is actually only productive about 30 percent of the time during their rotation week. The rest is context switching, checking dashboards, and worrying. Luna: And that's not even counting the recovery time after a bad week. It can take days to get back into flow. Lucas: Exactly. So the question is: do you want to continue subsidizing that hidden cost, or do you want to invest in a system that actually reduces it? The dedicated team model isn't the only answer, but it's a proven one. Luna: If today's conversation gave you something usable, a handful of listeners chip in monthly through buy me a coffee dot com slash fexingo, and that's literally what funds making this many of these. Lucas: Yeah, it keeps the show ad-free and lets us dig into topics like this without worrying about commercial breaks. Appreciate everyone who's already part of that. Luna: So back to the dedicated team model — what do you see as the biggest mistake people make when they try to implement it? Lucas: The biggest mistake is not giving the team enough authority to make changes. If the dedicated on-call team can't modify the system — can't push code changes, can't adjust alert thresholds — they're just triage operators. The whole point is that they should be empowered to fix the root causes. Luna: So they need to be part of the engineering team, not a separate support function. Lucas: Exactly. They should have the same access, the same deploy privileges, and the same responsibility as feature engineers. Otherwise you're just creating a tier-one support team with no ability to actually improve things. Luna: Another mistake I've seen is rotating people too quickly. If the tour is only a month, they barely learn the system before they're out. Lucas: Agreed. Three months seems to be the sweet spot. Long enough to develop expertise, short enough to avoid burnout. And you need overlap between teams — a week or two of shadowing before the new team takes over. Luna: What about the size of the dedicated team? Four seems like a good number for a large org, but what about smaller teams? Lucas: For a team of 20 to 30 engineers, you could have a two-person dedicated rotation with a backup from the larger team. The key is that the dedicated people are fully focused on reliability, not split between features and on-call. Luna: So the principle is: separate the roles clearly. Don't try to make engineers do deep feature work and be on the hook for incidents at the same time. Lucas: That's the core idea. And it seems obvious, but so many teams still run the old rotation because it's the default. I think in 2026, we have enough data to say that the default is broken. Luna: I agree. And the best engineering leaders I know are the ones willing to question the defaults, even when they're painful to change. Lucas: Exactly. If you're a CTO listening, I'd challenge you to audit your on-call setup. Calculate the real cost. Talk to your engineers about how it affects their work. You might be surprised at what you find. Luna: And if you do make a change, let us know how it goes. We love hearing from listeners who try things we discuss. Lucas: Absolutely. That's it for this episode of Tech Leadership with Fexingo. We'll be back next week with another conversation. Luna: See you then.