The Proliferation Problem: Why Engineering Teams Keep Spinning Up Clusters Instead of Optimizing the Ones They Have
Ask an engineering leader why their organization is running seventeen Kubernetes clusters when three would suffice, and you will rarely hear a satisfying answer. You will hear about a legacy migration that stalled. About a team that needed isolation for compliance reasons. About a proof-of-concept environment that somehow graduated to production. About a vendor requirement that seemed temporary. Each explanation is individually defensible. Collectively, they describe an infrastructure portfolio that has grown without direction and is now consuming budget, operational capacity, and engineering attention at a rate that no one explicitly approved.
This is cluster sprawl—and it is one of the most expensive problems that enterprise infrastructure teams consistently underestimate.
The Reasonable Decisions That Produce Unreasonable Outcomes
Cluster sprawl is rarely the product of bad engineering judgment. It is the product of good engineering judgment applied locally, without visibility into the broader infrastructure landscape.
Consider the sequence of events that typically produces a redundant cluster. A product team needs a new environment. The shared cluster that serves their organization is backlogged, subject to governance processes the team finds slow, or simply not configured in a way that supports their workload requirements. The path of least resistance is to provision a new cluster. The team has the cloud credentials, the tooling knowledge, and the budget authority to do it. They are up and running in hours rather than waiting weeks for a shared platform to accommodate them.
From that team's perspective, the decision was correct. From the organization's perspective, it added another environment to maintain, another security surface to monitor, another set of operational procedures to document, and another cost center to justify at the next budget review. Multiply this pattern across a dozen teams over three years and the result is an infrastructure portfolio that nobody designed and everyone depends on.
The Political Economy of Infrastructure Duplication
Beyond the technical friction that drives teams toward new clusters, there are organizational incentives that actively reward sprawl. In many enterprises, the team that builds something owns it—and ownership of infrastructure, however burdensome operationally, confers a degree of autonomy and control that teams are reluctant to surrender.
A team running its own cluster does not need to negotiate maintenance windows with a platform team. It does not need to wait for capacity to be allocated from a shared pool. It does not need to comply with configuration standards it did not write. The operational overhead is real, but so is the independence—and in environments where engineering velocity is the primary success metric, independence is worth paying for.
This dynamic is compounded by budget structures that make the true cost of a new cluster invisible at the team level. When compute costs are aggregated at the department or division level, the marginal cost of spinning up another cluster does not appear as a line item in any team's P&L. The team experiences the benefits of the new cluster immediately; the costs are distributed across the organization in ways that no individual team feels directly.
Tribal Knowledge Silos and the Consolidation Barrier
Even when enterprise leaders recognize the problem and commit to consolidation, the effort frequently stalls. The most common reason is not technical—it is epistemic. Individual clusters accumulate operational knowledge that lives in the heads of the engineers who built and maintain them. Configuration decisions that were made for reasons nobody documented. Dependency relationships that are not reflected in any CMDB. Tuning parameters that were set during an incident three years ago and never revisited.
When consolidation requires migrating workloads from a well-understood (if inefficient) cluster to a shared environment, the engineers responsible for that migration face a knowledge problem. They must understand the workload well enough to ensure it behaves correctly in the new environment, but the relevant knowledge is often tacit, partial, and distributed across team members who have other priorities.
This is why consolidation projects routinely exceed their timelines and budgets. The technical work of moving workloads is often straightforward. The work of understanding what the workloads actually require—and ensuring the destination environment can satisfy those requirements—is substantially harder.
The Metrics That Reveal the True Scale of the Problem
Organizations that want to address cluster sprawl need visibility that most of them do not currently have. The most important metrics are not the ones that infrastructure teams typically track.
Cluster-to-workload ratio is a useful starting point: how many active workloads does each cluster support? A portfolio with a high proportion of single-workload or dual-workload clusters is a strong indicator of sprawl driven by team-level provisioning rather than architectural planning.
Average cluster utilization—measured not at peak but across a rolling window—reveals how much of the provisioned capacity is actually being used. Clusters running at fifteen to twenty percent average utilization are common in sprawled environments and represent a significant and recoverable cost inefficiency.
Finally, operational cost per workload gives leadership a denominator that makes the cost of sprawl legible. A workload running on a dedicated cluster carries the full operational overhead of that cluster: patching, monitoring, security review, capacity management. The same workload on a well-managed shared cluster carries a fraction of that overhead. The difference, aggregated across a large portfolio, is often measured in millions of dollars annually.
What Consolidation Actually Requires
Addressing cluster sprawl requires changes at the organizational level, not just the technical one. Engineering teams will continue to provision independent clusters as long as the shared platform is slower, less flexible, or less capable than spinning up something new.
This means that platform teams must treat consolidation as a product problem. The shared cluster needs to be genuinely better—faster to provision, easier to operate, more capable of supporting diverse workloads—than the alternative. If the platform cannot make that case on its merits, governance mandates alone will produce resistance and workarounds rather than actual consolidation.
At the same time, budget structures need to be adjusted so that the true cost of independent clusters is visible at the team level. When teams bear the operational overhead of the clusters they build—in engineering time, in security review cycles, in on-call burden—the calculus of building versus consolidating shifts materially.
Cluster sprawl is, at its core, a story about misaligned incentives and invisible costs. The organizations that successfully reverse it are the ones that make those costs visible, reduce the friction of the shared alternative, and give teams a reason to consolidate that goes beyond compliance.