24/7 Cluster, Costs Down 60%: What We Did Differently

Written By

Nicola Petteta

Published

March 25, 2026

Kubernetes cluster nodes scaling up and down on AWS to reduce cloud costs

The right question at the right time

This isn't a blog post about Kubernetes. It's about asking the right question at the right time.

Nico, a Cloud Engineer at Renaiss, asked himself: "Why is this cluster still running 24/7?" And that simple question ended up saving Autoptic roughly 60% in operating costs on their Kubernetes cluster on AWS. But as in every good infrastructure story, the road from question to solution had its twists and turns. Here's what happened, what broke (or almost broke), and what Nico learned along the way.

The problem: when "it's always been like this" stops making sense

Autoptic had a Kubernetes cluster running 24/7 on AWS (EKS). At first, it made perfect sense. "It was useful for getting a constant picture of how the client's system behaved," Nico says. When you're starting out with a new client, you want to see everything: how the system behaves at different times of day, what its usage patterns are, where the bottlenecks show up.

But time went by. The cluster grew. What had started as a test environment with a handful of services was now running a wide variety of resources. And it was still on 24 hours a day, 7 days a week. "With the rapid growth, it was clear that at some point something would have to be done about the spending," Nico explains.

The problem was obvious: there was no longer any relevant information that justified keeping everything running around the clock.

The impact: when the problem is internal but still hurts

Unlike many infrastructure incidents, this one didn't directly affect end users. From their perspective, everything worked perfectly. The impact was internal: operating costs growing for no good reason. AWS resources being consumed outside business hours, when nobody really needed them.

"It's internal — nothing on the user side," Nico clarifies. "But in terms of services and the business, it delivered significant savings." There was also the opportunity to improve the architecture. The previous solution was simple: one big node running all the time. Moving to several nodes managed by Karpenter would provide "a glimpse of what a more 'real' solution looks like" — one better prepared to scale and adapt to actual demand.

The solution: Karpenter

The idea was clear: implement a smart way to remove resources outside business hours. Nico deployed Karpenter inside the cluster with jobs that scale resources up and down on a schedule. Essentially, going from "one big node always on" to "several nodes that Karpenter manages based on real demand."

"We implemented Karpenter inside the cluster with jobs that run both to scale up and to scale down resources outside business hours," he explains. Karpenter is an open-source autoscaler designed specifically for Kubernetes on AWS. Unlike other solutions, Karpenter can make more granular decisions about which instance types to use and when, optimizing costs and performance at the same time.

The result was striking: roughly a 60% reduction in the operating costs of Autoptic's cluster. But as with everything in infrastructure, the solution brought its own challenges and lessons.

What broke (or almost did)

"Did anything break afterwards?" we asked Nico. "Luckily, no," he answers. "But there's a fairly common problem with Karpenter that's important to know about." The problem is almost philosophical in its irony: Karpenter runs a controller that manages your cluster's nodes. And that controller also runs inside your cluster.

"If Karpenter removes the node where the controller lives for whatever reason, it's as if Karpenter deletes itself, and your cluster stays frozen as it is," Nico explains. It's the infrastructure equivalent of sawing off the branch you're sitting on. The system that decides which nodes to remove can, in theory, remove itself. Did it happen to Nico? No. But knowing it can happen changes how you design and implement the solution. As Werner Vogels, CTO of AWS, says: "In the world of software, everything breaks all the time." The difference lies in being prepared.

The lesson: always ship in small pieces

If there's one thing Nico took away from this project, it's an implementation philosophy he now applies to everything. "I learned to implement things in small pieces that keep working, instead of building it all at once," he says. "Once you have working pieces, it's much easier to see what breaks as you add something new." Also: test at a smaller scale before rolling out. Don't start with the whole production cluster. Begin with a controlled environment, verify that it works, and then scale.

The constant trade-off: make it work or make it perfect?

We asked Nico whether at any point he had to choose between "make it perfect" and "make it work." "Yes," he answers without hesitation. "Since things were implemented piece by piece, it was mostly about making things work until they were done right." That tension between speed and perfection is a constant in infrastructure. And Nico's answer is pragmatic: make it work first, make it perfect later. But always with an eye on improving. Because in cloud infrastructure, things break. The difference lies in being prepared, shipping in small pieces, and learning along the way. Nico's question — "why is this cluster still running 24/7?" — is the one that rarely gets asked when the team is focused on keeping everything running. Reviewing your architecture with that mindset and finding the 60% of cost you don't need is a big part of what we do in cloud cost optimization. If your AWS bill is growing faster than your business, there's probably a question like that waiting for someone to ask it.

Want to optimize your cloud infrastructure?

At Renaiss we help companies solve real business problems through technology. From cost optimization to scalable cloud architecture.
If you have a cluster running 24/7 and you're wondering whether it really makes sense, let's talk. We can assess your current infrastructure and find optimization opportunities.