Latest / Tech Leadership with Fexingo: Engineering Managers, CTOs, and Technical Leadership Conversations / How a CTO Uses DORA Metrics to Measure Engineering Performance
Transcript
- Lucas: There's a moment every CTO hits, usually about eighteen months into scaling a team from twenty engineers to forty or fifty, where you realize you have no idea whether you're actually getting faster. Luna: Right. You're shipping features, you're doing retrospectives, everyone feels busy — but 'busy' isn't a metric. Lucas: Exactly. And that's where the DORA metrics come in. DORA stands for DevOps Research and Assessment, which started as a Google research project. They boiled software delivery performance down to four numbers: deployment frequency, lead time for changes, change failure rate, and time to restore service. Luna: I've seen these in a bunch of engineering blogs, but I've never actually watched a CTO implement them from scratch. What does that look like on the ground? Lucas: Let me give you a concrete case. I was talking to a CTO at a mid-market SaaS company — about forty engineers, B2B analytics platform. They were doing monthly releases, change failure rate was around fifteen percent, and mean time to restore was something like eight hours. That means one in seven deployments broke something, and when it did, it took a full day to fix. Luna: That's painful. And I'm guessing the team felt it, even if they didn't have the numbers to prove it. Lucas: They absolutely did. The CTO told me the engineers were exhausted. But here's the thing — they didn't know their own baseline. So the first step was just measuring. They instrumented their CI/CD pipeline to track deployment frequency and lead time. For change failure rate, they tagged any incident that was linked to a deployment. And time to restore came from their incident management tool. Luna: So the act of measuring itself was the intervention. Did they make a big announcement about it? Lucas: They did the opposite. The CTO introduced the metrics in a quiet all-hands, framed entirely as a diagnostic. He said, 'We're going to collect these numbers for a quarter, no targets, no judgment. We just want to see where we are.' And the data was rough. Deployment frequency: once every three weeks. Lead time: about twelve days from commit to production. Luna: Which is basically batch and queue. Not continuous delivery at all. Lucas: Right. And that slow cadence actually made the change failure rate worse, because each deployment was a huge batch of changes. Hard to isolate what caused the failure. So the CTO's first move was to target lead time. They broke their monorepo into smaller deployable units, introduced feature flags so they could decouple deploy from release, and moved to a trunk-based development model. Luna: That's a lot of change. How did the engineers react? Lucas: Mixed. Some loved it — the senior engineers had been pushing for trunk-based development for years. Others were nervous about the feature flags, worried they'd create complexity. But the CTO ran a small pilot with one team for four weeks. That team's lead time dropped from twelve days to two days, and their deployment frequency went from once every three weeks to twice a week. Luna: And the change failure rate? Lucas: Initially it went up. Because they were deploying more often, and not all the automation was in place. But after about six weeks, it started coming down. By the end of six months, their change failure rate was under three percent, and mean time to restore was under thirty minutes. Luna: That's a dramatic shift. Six months from monthly releases with a fifteen percent failure rate to multiple deploys a week with near-zero failure. Lucas: Honestly, if today's tech conversation gave you something usable, it might be worth the price of a coffee. If it was, you can find us at buy me a coffee dot com slash fexingo. That's all lowercase, dot com slash fexingo. No pressure — just a way to keep this ad-free. Luna: Yeah, and we mean it. No perks, no tiers. Just a link if the episode was valuable. Lucas: So back to the DORA story. One of the things this CTO did that I thought was smart was he never used the metrics to evaluate individuals. No one got a bad performance review because their personal change failure rate was high. The metrics were always team-level, and they were always used to ask 'what's in our way?' not 'who messed up?' Luna: That's critical. Because if you start using DORA metrics as a stick, people will game them. You'll get tiny, meaningless deploys just to inflate the frequency number. Lucas: Exactly. Goodhart's law in action. The CTO told me he saw that happen at a previous company. A team started deploying empty commits just to boost their deployment frequency. So he built in a quality gate: each deploy had to include at least one user-facing change or bug fix. And they tracked the median, not the average, for lead time, to avoid outliers skewing the picture. Luna: Median is a smart choice. Averages can hide a lot. If most deploys take two hours but one takes two weeks, the average looks terrible even if the system is mostly fast. Lucas: Right. And they also paired DORA with a second metric: customer-facing uptime. Because it's possible to have great DORA scores but still have a lousy product experience if your architecture is fragile. So they tracked whether customers actually noticed improvements. Luna: Did they see a correlation? Better DORA scores leading to happier customers? Lucas: They did, but not immediately. In the first two months, customer satisfaction actually dipped slightly — because the team was shipping more frequently, and some of those changes introduced minor bugs. But by month four, the trend reversed. The team was able to respond to customer feedback much faster, so issues got fixed within hours instead of weeks. Net Promoter Score went up by twelve points over the year. Luna: So the key was persistence. They didn't panic when the metrics wobbled early on. Lucas: Exactly. And that's the hardest part for most CTOs — staying the course when the numbers get worse before they get better. But the CTO had a mantra: 'Measure less, but measure the right things.' He only tracked those four DORA metrics plus uptime and NPS. That was it. Luna: I like that. Most engineering orgs I see are drowning in dashboards. Cycle time, velocity, story points, bug counts, code coverage... it's noise. Lucas: It is. And the DORA research actually backs that up. The teams that improved the most were the ones that focused on a small set of outcomes, not a large set of outputs. So if you're a CTO listening and you're thinking about starting, my advice is: just pick deployment frequency and change failure rate. Two metrics. Measure them for a month. See what they tell you. Luna: And use them to start conversations, not end them. Lucas: Exactly. One team might have a high change failure rate because their test coverage is low. Another might have low deployment frequency because their code review process is a bottleneck. The same metric points to two different root causes. The numbers are just the starting point. Luna: I want to circle back to something you mentioned — the median lead time. How did they actually measure that? Lucas: They instrumented their CI/CD pipeline to capture the timestamp when a commit was first pushed to the main branch, and then the timestamp when that commit was deployed to production. The difference is the lead time. They calculated the median across all deploys in a given week. That gave them a stable trend. Luna: And for time to restore? That's often the trickiest one to instrument. Lucas: It is, because you have to define 'restore.' They used the time from when a deployment-related incident was declared — that could be a page or a Slack alert — to when the service was back to normal, measured by their monitoring. They excluded any incident that wasn't linked to a deploy, to keep the metric focused on delivery quality. Luna: That's a good distinction. Otherwise you're measuring overall system reliability, not the quality of your deployment process. Lucas: Right. And the CTO made sure the team knew the difference. He published a simple one-pager: 'What DORA means for us.' It defined each metric, explained why it mattered, and explicitly said 'these are team metrics, not individual metrics.' That transparency built trust. Luna: What about the skeptics? Were there engineers who thought this was just another management fad? Lucas: A few. But he won them over by letting them see the data. He showed them the trend over time — the lead time graph sloping down, the deployment frequency sloping up. And he pointed to specific incidents that were caught faster because the team was deploying more often. One senior engineer told me, 'I was against it until I saw our time to restore drop from eight hours to forty minutes. That changed my mind.' Luna: That's the best kind of proof. Real numbers, real improvement. Lucas: And the CTO himself learned something. He told me he initially thought deployment frequency was the most important metric. But after six months, he realized that time to restore was actually the leading indicator. If you can fix failures fast, you can afford to deploy more often, which reduces batch size, which reduces failure risk. It's a virtuous cycle that starts with recovery speed. Luna: That's a great insight. So for any CTO starting out, maybe lead with that — invest in your incident response and rollback capabilities first. Lucas: I think that's exactly right. The DORA metrics are a system, not a checklist. And the best way to use them is to pick one, improve it, and watch the others follow. Over a year, that is what happened at this company.