Latest / Tech Leadership with Fexingo: Engineering Managers, CTOs, and Technical Leadership Conversations / How One CTO Cut Cloud Costs by Rethinking Commitments
Transcript
- Lucas: So there's this CTO at a mid-stage SaaS company — about 200 employees, running on AWS — who told me his cloud bill was growing thirty percent year-over-year, and he couldn't point to a single new feature driving it. Luna: That's the classic 'cloud cost drift' story. No single decision, just death by a thousand small commits. Lucas: Exactly. And he knew the usual playbook — turn off idle resources, pick smaller instance types — but he'd already done that. The real lever he hadn't touched was how he committed to cloud spend. Luna: Commitment-based discounts. Reserved instances, Savings Plans, that stuff. Lucas: Right. And honestly, if this conversation saves you even five percent on your next cloud bill, that's the kind of thing that makes the show worth it. If today's tech conversation gave you something usable, you can find the link at buy me a coffee dot com slash fexingo. Luna: Yeah, we keep the show ad-free, and listener support is what makes that possible. Appreciate anyone who chips in. Lucas: So back to this CTO. He had about five hundred thousand dollars a year in compute spend, with roughly sixty percent covered by one-year all-upfront reserved instances. The rest was on-demand. Luna: Pretty typical mix. What did he change? Lucas: He moved to a portfolio of convertible reserved instances and Savings Plans. But the key wasn't the product — it was the process. He started a rolling three-month forecast of compute demand by workload. Luna: So instead of guessing at the beginning of the year, he refreshed the forecast every quarter. Lucas: Exactly. And that allowed him to do three things. First, he converted his underutilized one-year reservations into convertible ones — same discount, but now he could change instance families if his workload shifted. Luna: That's smart. A lot of teams lock themselves into a specific instance type and then can't pivot. Lucas: Second, he layered a three-year partial upfront Savings Plan on top of his reservations. That gave him an additional five percent discount on his baseline spend. And third, he rightsized every instance that had been running for more than six months based on actual CPU and memory utilization. Luna: Rightsizing is tricky though. If you downsize and then hit a spike, you degrade performance. Lucas: He used a buffer. He targeted average utilization between sixty and seventy percent, not the typical ninety percent that a lot of optimization tools suggest. That gave him headroom for spikes without paying for idle capacity. Luna: And what was the net effect? Lucas: Over twelve months, he reduced total compute cost by forty percent — about two hundred thousand dollars. Roughly half came from the commitment portfolio changes, the other half from rightsizing. Luna: That's a huge number. But does this work for startups with spiky traffic? Like a company that sees two times load on certain days? Lucas: It's harder. He told me that for variable workloads, he actually kept some instances on-demand and used spot instances for batch jobs. The key is knowing your 'commitment delta' — the gap between what you reserved and what you actually use. Luna: Commitment delta. I like that. So what's a healthy delta? Lucas: He said under ten percent is good. If your delta is over thirty percent, you're leaving money on the table — either over-reserved or under-utilized. He tracks it monthly on a dashboard. Luna: That's a concrete metric every engineering leader can audit this quarter. Just pull your reserved instance coverage report and calculate your unused hours. Lucas: Exactly. And it doesn't require a big tooling investment. AWS Cost Explorer, Azure Cost Management, GCP's Committed Use Discount reports — they all give you this data. Luna: So the real lesson isn't about specific services. It's about having a dynamic commitment strategy tied to a rolling forecast, not a static annual purchase. Lucas: That's the core. And he also mentioned one more thing: he involved his finance team in the forecasting. They had a monthly meeting where engineering presented the three-month compute forecast, and finance incorporated it into the budget. Luna: So it became a cross-functional process. That probably helped with buy-in. Lucas: Huge. Before, engineering would buy reservations reactively. After, they had a cadence. And the finance team actually started asking questions like 'why is our forecast for database instances growing faster than our user base?' Luna: That's the kind of conversation most companies never have. They just pay the bill. Lucas: Right. And the CTO said the biggest win wasn't even the money — it was the cultural shift. Engineering started thinking about cost efficiency as a design constraint, not an afterthought. Luna: So what's the one thing a listener should do Monday morning? Lucas: Calculate your commitment delta. Go into your cloud console, pull your reservation coverage report for the last thirty days, and see what percentage of your reserved instances went unused. If it's over twenty percent, start a conversation about converting those reservations or letting them expire. Luna: And if it's under ten, you're in good shape. But still check your Savings Plans — they might give you an extra five percent. Lucas: Yeah. And if you're not using any commitments at all, you're probably overpaying by fifteen to twenty percent. That's a quick win. Luna: It's funny — cloud cost optimization sounds like a boring ops topic, but it's actually a leadership lever. It forces you to understand your architecture, your growth, and your business priorities. Lucas: Totally. And it's a concrete way to show the board that engineering is thinking about the bottom line. That matters when you're asking for headcount next quarter. Luna: Alright, next time we should talk about how to get finance to actually trust engineering forecasts. Lucas: That's a whole episode. And we'll do it.