Resources
OptOps Blog
Insights, playbooks, and field notes on infrastructure cost optimization — Kubernetes, HPC, and GPU compute — written by the team building OptOps.ai.
One estate, two schedulers: optimizing Kubernetes and Slurm together.
Organisations rarely choose between Slurm and Kubernetes — they end up with both, and nobody can answer a simple question about total cost across the two. Why the waste patterns differ, why gang scheduling matters, and how to build one view across the estate.
the lit ones do the work. you pay for all of them.
Read articleTechnical · GPU Observability
Measuring GPU idle waste: what DCGM tells you that nvidia-smi does not.
The utilization number almost everyone tracks is probably the most misread metric in accelerated computing. What it actually measures, the DCGM fields that answer the real question, and how to make per-job idle waste calculable.
Sovereign Compute · Policy & Infrastructure
Sovereign compute economics: why national AI infrastructure cannot use cloud cost tools.
National AI programmes are building compute at extraordinary pace — and none of it can use the tools that manage cloud cost. Four assumptions that break in sovereign environments, and what optimisation has to look like instead.
HPC · Cluster Efficiency
Why researchers over-request GPUs, and what to do about it.
Every HPC centre sends the same message asking people to request resources accurately. Nothing changes, and six months later they send it again. The message fails for a reason worth understanding — and the fixes that work change defaults and visibility, not policy.
HPC · Pillar Post
The HPC utilization problem: why a cluster at 95% can still be half idle.
Supercomputing centres report utilisation in the eighties and nineties, and the numbers are honest. They are also measuring the wrong thing. Where GPU waste actually hides, why nobody sees it, and the five metrics that describe reality rather than bookkeeping.
Buyer's Guide · 2026 Edition
A field guide to Kubernetes cost optimization tools. 2026 edition.
The market has consolidated into three categories, each built on a different philosophy. This guide covers what each one does, who builds in it, and which environments fit. Written by a vendor in this space, but kept as straight as we can manage.
Industry Analysis · Opinion
Why your cloud providers don't offer real cost optimizers.
Hyperscalers ship billing dashboards, advisor services, and recommender tools. They don't ship platforms that meaningfully cut your bill. There is a structural reason, and it isn't a conspiracy. It's basic economics.
BFSI · Compliance Playbook
Kubernetes cost optimization for BFSI. A compliance-aware playbook.
Banks and financial services run some of the most sophisticated Kubernetes infrastructure in the world, and pay some of the worst Kubernetes bills. The reason isn't engineering skill. It's that most optimisation tools were built for environments where compliance isn't the lead constraint.
Opinion · Engineering Culture
The real reason your engineering team ignores cost optimization recommendations.
Not laziness, not lack of awareness, and not bad attitude. The cause is the incentive structure your organisation built, plus recommendations that don't match how engineers evaluate risk. Inside, the honest version written by someone who has been on both sides.
Technical · Pod Rightsizing
Pod rightsizing done right. Why VPA isn't enough, and what to use instead.
The Kubernetes Vertical Pod Autoscaler is built into Kubernetes for a reason. It's also why almost nobody uses it in production. A technical breakdown of what VPA does, where it fails, and what to use when you need pod rightsizing that works.
Government · Public Sector
Kubernetes cost optimization in government and public sector. What's different.
Government Kubernetes deployments don't operate under the same constraints as commercial infrastructure. The optimisation playbook that works for a Series B SaaS company can fail in a public-sector context for reasons that aren't immediately obvious. What changes, and what teams running cloud at scale in government need.
Buyer's Guide · Decision Framework
Advisory vs. autonomous Kubernetes optimization. Which approach fits your team?
Every Kubernetes cost optimization platform on the market is built on one of two opposing philosophies. Pick the wrong one and the tool either gets blocked in security review or quietly stops being used. Inside, a framework for choosing right.
FinOps · Pain Piece
Why FinOps dashboards don't save you money. And what does.
Your FinOps dashboard is doing exactly what it was built to do. Showing you where Kubernetes waste lives. The reason your bill isn't going down has nothing to do with the dashboard. It's the part the dashboard was never designed to handle.
Kubernetes Cost · Pillar Post
The hidden cost of Kubernetes over-provisioning. Why 60% of your cluster spend is probably waste.
Half of what you're paying for Kubernetes is sitting idle. This piece walks through how that happens, why the obvious fixes fall short, and what teams running production at scale do about it.
Educational · Pillar Piece
Reading the Kubernetes cost stack. Pods, nodes, and cloud commitments.
Kubernetes cost lives in three distinct layers, each with its own waste pattern, its own owner, and its own optimisation approach. Most teams attack one layer at a time and miss the compounding effects. A way to read the whole stack.
Founder Story · Origin
The founder pivot. How watching our cloud bill blow up built a company.
A blindside AWS bill. A co-founder who saw what nobody else was seeing. A 60% drop that nobody believed was real until it was. The story of how OptOps.ai started, written by the people it happened to.

