Resources

OptOps Blog

Insights, playbooks, and field notes on infrastructure cost optimization — Kubernetes, HPC, and GPU compute — written by the team building OptOps.ai.

Featured · Technical · Multi-Scheduler

One estate, two schedulers: optimizing Kubernetes and Slurm together.

Organisations rarely choose between Slurm and Kubernetes — they end up with both, and nobody can answer a simple question about total cost across the two. Why the waste patterns differ, why gang scheduling matters, and how to build one view across the estate.

By Ashutosh Dubey, Co-founder & CTO·August 11, 2026·9 min read
Read article
All posts

Technical · GPU Observability

Measuring GPU idle waste: what DCGM tells you that nvidia-smi does not.

The utilization number almost everyone tracks is probably the most misread metric in accelerated computing. What it actually measures, the DCGM fields that answer the real question, and how to make per-job idle waste calculable.

August 4, 2026 · 8 min read

Sovereign Compute · Policy & Infrastructure

Sovereign compute economics: why national AI infrastructure cannot use cloud cost tools.

National AI programmes are building compute at extraordinary pace — and none of it can use the tools that manage cloud cost. Four assumptions that break in sovereign environments, and what optimisation has to look like instead.

July 28, 2026 · 9 min read

HPC · Cluster Efficiency

Why researchers over-request GPUs, and what to do about it.

Every HPC centre sends the same message asking people to request resources accurately. Nothing changes, and six months later they send it again. The message fails for a reason worth understanding — and the fixes that work change defaults and visibility, not policy.

July 21, 2026 · 7 min read

HPC · Pillar Post

The HPC utilization problem: why a cluster at 95% can still be half idle.

Supercomputing centres report utilisation in the eighties and nineties, and the numbers are honest. They are also measuring the wrong thing. Where GPU waste actually hides, why nobody sees it, and the five metrics that describe reality rather than bookkeeping.

July 14, 2026 · 9 min read

Buyer's Guide · 2026 Edition

A field guide to Kubernetes cost optimization tools. 2026 edition.

The market has consolidated into three categories, each built on a different philosophy. This guide covers what each one does, who builds in it, and which environments fit. Written by a vendor in this space, but kept as straight as we can manage.

May 2026 · 12 min read

Industry Analysis · Opinion

Why your cloud providers don't offer real cost optimizers.

Hyperscalers ship billing dashboards, advisor services, and recommender tools. They don't ship platforms that meaningfully cut your bill. There is a structural reason, and it isn't a conspiracy. It's basic economics.

May 2026 · 9 min read

BFSI · Compliance Playbook

Kubernetes cost optimization for BFSI. A compliance-aware playbook.

Banks and financial services run some of the most sophisticated Kubernetes infrastructure in the world, and pay some of the worst Kubernetes bills. The reason isn't engineering skill. It's that most optimisation tools were built for environments where compliance isn't the lead constraint.

April 2026 · 9 min read

Opinion · Engineering Culture

The real reason your engineering team ignores cost optimization recommendations.

Not laziness, not lack of awareness, and not bad attitude. The cause is the incentive structure your organisation built, plus recommendations that don't match how engineers evaluate risk. Inside, the honest version written by someone who has been on both sides.

April 2026 · 7 min read

Technical · Pod Rightsizing

Pod rightsizing done right. Why VPA isn't enough, and what to use instead.

The Kubernetes Vertical Pod Autoscaler is built into Kubernetes for a reason. It's also why almost nobody uses it in production. A technical breakdown of what VPA does, where it fails, and what to use when you need pod rightsizing that works.

April 2026 · 9 min read

Government · Public Sector

Kubernetes cost optimization in government and public sector. What's different.

Government Kubernetes deployments don't operate under the same constraints as commercial infrastructure. The optimisation playbook that works for a Series B SaaS company can fail in a public-sector context for reasons that aren't immediately obvious. What changes, and what teams running cloud at scale in government need.

April 2026 · 9 min read

Buyer's Guide · Decision Framework

Advisory vs. autonomous Kubernetes optimization. Which approach fits your team?

Every Kubernetes cost optimization platform on the market is built on one of two opposing philosophies. Pick the wrong one and the tool either gets blocked in security review or quietly stops being used. Inside, a framework for choosing right.

March 2026 · 8 min read

FinOps · Pain Piece

Why FinOps dashboards don't save you money. And what does.

Your FinOps dashboard is doing exactly what it was built to do. Showing you where Kubernetes waste lives. The reason your bill isn't going down has nothing to do with the dashboard. It's the part the dashboard was never designed to handle.

March 2026 · 7 min read

Kubernetes Cost · Pillar Post

The hidden cost of Kubernetes over-provisioning. Why 60% of your cluster spend is probably waste.

Half of what you're paying for Kubernetes is sitting idle. This piece walks through how that happens, why the obvious fixes fall short, and what teams running production at scale do about it.

February 2026 · 9 min read

Educational · Pillar Piece

Reading the Kubernetes cost stack. Pods, nodes, and cloud commitments.

Kubernetes cost lives in three distinct layers, each with its own waste pattern, its own owner, and its own optimisation approach. Most teams attack one layer at a time and miss the compounding effects. A way to read the whole stack.

February 2026 · 10 min read

Founder Story · Origin

The founder pivot. How watching our cloud bill blow up built a company.

A blindside AWS bill. A co-founder who saw what nobody else was seeing. A 60% drop that nobody believed was real until it was. The story of how OptOps.ai started, written by the people it happened to.

January 2026 · 8 min read
Calculate ROI