Latest / Tech Leadership with Fexingo: Engineering Managers, CTOs, and Technical Leadership Conversations / How One Engineer Changed Our Entire Deployment Pipeline
Transcript
- Lucas: So there's this story from a mid-size SaaS company — about 80 engineers — where one senior engineer essentially rewrote their entire deployment pipeline in six weeks. Single-handedly. Luna: One person? That sounds like either a hero story or a recipe for disaster. Lucas: Right, and that tension is exactly what I want to dig into today. Before this engineer, their deployments were a two-hour manual process. You had to be on a Zoom call, follow a 37-step checklist, and if anything went wrong, the rollback took 45 minutes. They deployed once a week — Thursday afternoons — and if you missed the window, your code sat for another week. Luna: I've lived that. The Thursday afternoon dread is real. Lucas: Exactly. So this engineer, let's call her Priya, proposed a fully automated pipeline. She said she could build it in two months. And the CTO said yes — but with a condition: she had to keep doing her normal feature work. She'd work on the pipeline in whatever time she had left. Luna: Which basically means evenings and weekends, right? That's a fast track to burnout. Lucas: That's what happened. She burned herself out in about three weeks. Then she went to the CTO and said, 'Look, I need dedicated time. Give me six weeks, no feature work, and I'll deliver.' And he agreed. Luna: That's a bold ask. Most managers would say no. Lucas: Right. And here's where it gets interesting. The CTO told me later that the reason he said yes was because he'd seen Priya's design doc — she'd mapped out every error state, every rollback scenario, every integration point. It was 20 pages of precise architecture. He trusted her. Luna: That trust is rare. And honestly, if you have an engineer who can deliver something that transformational, you owe it to them to protect their time. Lucas: That's actually a great segue. Because at Fexingo Business, we try to protect our listeners' time by keeping the show ad-free and focused. If today's conversation gave you something useful — maybe a framework for thinking about who gets to own critical infrastructure — and you'd like to support that, you can buy us a coffee. It's buy me a coffee dot com slash fexingo. All lowercase. And it genuinely helps keep this podcast independent. Luna: Yeah, no pressure at all. If you find value, great. If not, the next episode is just as good. Now, back to Priya — she delivered the pipeline in five and a half weeks, right? Lucas: Six weeks on the nose. She automated everything — build, test, deploy, rollback. Deployments went from weekly to seven times a day. Rollback time dropped from 45 minutes to under two minutes. And the on-call team went from getting paged three times per deployment to maybe once a month. Luna: Those are insane numbers. But I'm guessing there's a downside. Lucas: There is. Priya became the single point of failure. No one else understood the pipeline deeply. When she took vacation, deployments slowed down. When she left a year later, the pipeline started degrading. Within three months, they were back to manual deployments twice a week. Luna: So the classic bus-factor problem. You let one person own the critical path without building redundancy. Lucas: Exactly. And this is the lesson I want to talk about. When you give an engineer that kind of ownership, you have to also build in knowledge transfer. The CTO told me later that his biggest regret wasn't letting Priya own the pipeline — it was not requiring her to document and pair-program with a second engineer during the last two weeks. Luna: So what did they do when she left? Lucas: They hired a platform engineer to rebuild it from scratch, but this time with a team of two. They also created a rotating ownership model — every quarter, a different engineer from the platform team leads the pipeline for a month. That way, the knowledge spreads naturally. Luna: Rotating ownership. I like that. It solves both the bus-factor problem and gives junior engineers a chance to learn critical systems. Lucas: Right. And the second pipeline took four months to build — longer than Priya's — but it was more modular. They broke it into stages: build, test, deploy, canary, full rollout. Each stage had its own owner. So if one broke, you didn't lose the whole pipeline. Luna: That modularity is a big deal. It also makes it easier to experiment — you can swap out the canary step, for instance, without touching the rest. Lucas: Exactly. And the deployment frequency actually surpassed Priya's numbers — they hit twelve deployments a day at peak. Because they could parallelize. Multiple teams could ship at the same time. Luna: So the lesson isn't 'don't let one engineer own critical infrastructure.' It's 'if you do, plan for their departure from day one.' Lucas: That's a great way to put it. And I think there's another layer: the cultural shift. Before Priya, the team saw deployment as a painful gate. After her pipeline, they saw it as a routine step. But when she left and the pipeline degraded, that trust was broken. It took months to rebuild. Luna: Trust in the system is fragile. If you've ever had a deployment break production, you know how fast people start avoiding the pipeline again. Lucas: Right. So the CTO also invested in automated testing and monitoring specifically for the pipeline itself. Not just the app, but the pipeline. They added a dashboard showing pipeline health — green, yellow, red. If it turned yellow, someone was paged within five minutes. Luna: Monitoring the pipeline like it's a production service. That's smart. Lucas: And they wrote an architectural decision record documenting every major decision — why they chose a particular tool, why they structured stages a certain way. That ADR became the onboarding document for any new platform engineer. Luna: I think we've talked about ADRs before, but this is a perfect example of why they matter. Without that document, the second team would have had to reverse-engineer everything. Lucas: Exactly. So when I talk to CTOs now about building critical infrastructure, I always ask: 'Who will understand this system if the person who built it wins the lottery tomorrow?' If they hesitate, that's a red flag. Luna: And if they say 'no one,' you know you have work to do. Lucas: Right. But I also don't want to discourage teams from letting a single engineer take ownership of a big project. Priya's first pipeline was genuinely transformational. The mistake was in the handoff, not the ownership. Luna: So you'd still let a senior engineer go solo on a critical project today? Lucas: I would, but with three conditions. One: they have to document as they go — ADRs, runbooks, architecture diagrams. Two: they have to pair with at least one other engineer for the last 20 percent of the project. Three: after delivery, they need a two-week 'knowledge transfer sprint' where they teach a rotation of engineers how it works. Luna: That last one sounds like it could slow down velocity, but I think it actually accelerates long-term resilience. Lucas: Exactly. And the CTO I talked to said that after the second pipeline, they made knowledge transfer a formal part of any critical infrastructure project. It's a line item in the project plan now. Luna: So if you're a engineering manager listening, what's the one thing you should take away from Priya's story? Lucas: I'd say: before you give a single engineer ownership of a critical system, ask yourself what the plan is for when they leave. And if you don't have a good answer, start building one before they start coding. Luna: That's a solid takeaway. And maybe the next time you see an engineer burning out on a side project like that, you step in earlier. Lucas: Absolutely. Priya's burnout was avoidable. A good manager would have noticed the late-night commits and pulled her aside. But that's a topic for another episode. Luna: Fair enough. Thanks, Lucas. Lucas: Thanks, Luna. And thanks to everyone listening. We'll be back tomorrow.