There is a category of compute that the cloud cost management industry has quietly ignored.
It is growing quickly. India’s IndiaAI Mission was approved by the Union Cabinet in March 2024 with a total outlay of ₹10,371.92 crore, roughly 1.25 billion dollars over five years, and compute capacity is its largest single pillar. As of March 2026 the government reported more than 38,000 GPUs onboarded through the mission’s common compute facility, a figure officially stated as 38,231. Separately, the National Supercomputing Mission had deployed 37 supercomputers totalling 40 petaflops as of December 2025.
Comparable programmes are running across the Gulf, and data localisation requirements are driving similar builds across Africa and Southeast Asia.
None of this infrastructure can use the tools that manage cloud cost, and the reason is not technical readiness. It is architectural.
What sovereign compute actually means
The term gets used loosely, so it is worth being precise.
Sovereign compute is infrastructure where the physical location of the hardware, the jurisdiction governing the data, and the chain of control over operations are all deliberately constrained. Typically this means the hardware sits in-country, the data never leaves national boundaries, and no foreign entity holds operational control over the systems.
In practice it covers national supercomputing facilities, government AI compute programmes, defence and intelligence infrastructure, central bank and financial regulator systems, and increasingly the on-premise estates that regulated enterprises build because their data cannot go anywhere else.
The common thread is not secrecy. It is that the normal assumptions of cloud operations do not hold.
Four assumptions that break
That telemetry can leave the environment
Almost every cloud cost platform works by shipping metrics to a vendor’s infrastructure, analysing them there, and returning recommendations through a hosted console.
For sovereign environments this is frequently a non-starter, and the objection is not about how sensitive utilisation metrics look in isolation. Aggregate infrastructure telemetry from a national facility reveals capacity, workload patterns, procurement cycles and operational tempo. That is strategically meaningful regardless of how mundane any individual data point appears.
Anything requiring an outbound connection to a vendor’s cloud is excluded before the technical evaluation begins.
That the environment is connected at all
A significant share of this infrastructure is air-gapped or sits behind one-way data diodes. There is no outbound path to license-check, no telemetry pipeline, no automatic updates, and no possibility of a hosted control plane.
Software that assumes connectivity does not degrade gracefully in these environments. It simply does not run.
That waste shows up as a bill
This is the deepest difference and the one that most changes what optimisation has to mean.
Cloud cost optimisation is built on a feedback loop. Consumption produces an invoice, the invoice creates pressure, the pressure funds the work. Sovereign infrastructure is capital expenditure. The hardware was procured, commissioned and depreciated. Whether it runs at 40% or 90% effective efficiency, this month costs exactly the same.
So there is no invoice to react to, and every tool built around bill analysis has nothing to analyse.
The cost is real, just denominated differently. It appears as queue time for researchers who cannot get allocation, as capability a national programme cannot deliver on hardware it already owns, as power drawn by nodes doing little useful work, and as procurement cycles arriving earlier than they needed to. For a mission with a fixed multi-year budget, that last one is the expensive currency. Capacity recovered from existing systems is capacity that does not have to be bought again.
That write access is a technical question
In commercial cloud environments, granting a management platform permission to change infrastructure is a risk assessment. In sovereign environments it is frequently a policy prohibition, and policy does not yield to arguments about how reliable the automation is.
Separation of duties requirements in most government change-management frameworks are explicit that the party proposing a change cannot be the party executing it without authorisation. An autonomous optimiser collapses those roles by design. During an audit, attributing a production change to an optimiser is not an acceptable answer.
What optimisation has to look like instead
Take those four constraints seriously and the shape of a workable approach follows directly.
Analysis runs where the data is. Telemetry is collected, processed and turned into recommendations inside the customer’s environment. Nothing about workload behaviour crosses the boundary. This has to be decided early, because retrofitting it into a hosted platform is close to a rewrite.
It works disconnected. No outbound calls as a condition of functioning. No hosted control plane holding state the cluster depends on.
Read-only is the default posture. The system observes and recommends. Humans decide. Where changes are applied, the permission is scoped to that specific approved action, time-bound and revocable, with the audit trail landing in systems the operator already runs rather than in a vendor console.
Value is expressed in capacity, not currency. A recommendation that says a change saves a certain number of dollars per month is meaningless when the hardware is already bought. The useful framing is node hours recovered, queue time reduced, additional jobs served on the same installed base. That is what a facility director reports upward, and it is what determines whether the next procurement can be deferred.
It handles heterogeneous hardware. Sovereign estates accumulate. Systems arrive in procurement waves years apart, mixing accelerator generations, interconnects and schedulers. Anything assuming a uniform fleet fails on contact.
The thing nobody wants to say out loud
Facilities of this kind report utilisation figures in the eighties and nineties, and those numbers are honest. They are also, almost always, measuring allocation rather than consumption.
India’s own reporting illustrates how soft these numbers can be. In August 2025 the government described NSM systems as running at over 85% capacity with many exceeding 95%. In December 2025, with the same 37 systems and the same 40 petaflops installed, the reported figure was over 81% with few exceeding 95%. No official source defines what the percentage measures.
That distinction matters because a job requesting eight GPUs and saturating two still registers as eight GPUs occupied, and a twelve-hour allocation finishing its real work in four still registers twelve hours used. Research analysing more than 118,000 jobs on the Perlmutter system at NERSC found mean peak GPU memory utilisation of 28.64%, with over half of workloads showing consistently low memory utilisation, even though compute utilisation was generally well balanced.
This creates an uncomfortable dynamic for facility operators. Reporting high utilisation is how expansion funding gets justified. Investigating effective utilisation risks producing a number that complicates that case.
These are different metrics answering different questions, and both are worth having. A facility that can demonstrate it recovered 20% more throughput from existing hardware is making a stronger argument for expansion, not a weaker one, because it is showing stewardship of capital already committed. Reviewers respond to that. The alternative, where a reviewer discovers the gap independently, is considerably worse.
What this means if you operate one of these facilities
Three things worth doing regardless of what tooling you eventually adopt.
Measure consumption, not just allocation. Join scheduler records with device telemetry and produce a requested-against-consumed figure per job. Most facilities have both data sources already and have never joined them.
Establish the access position before evaluating anything. Decide what a third-party system would be permitted to see, whether anything may leave the environment, and whether write access is available at all. Answering this first eliminates whole categories of tooling before anyone spends a quarter on a pilot that was never deployable.
Express efficiency in mission terms. Node hours recovered and jobs served translate into research output and deferred procurement. Currency saved does not, when nobody is paying a monthly bill.
Why this matters beyond any single facility
National compute programmes are funded on the argument that domestic AI and scientific capability depends on domestic infrastructure. That argument is sound. It also carries an obligation, because a mission with a fixed budget that runs its hardware inefficiently delivers less capability than the funding was meant to buy.
Efficiency in this context is not a cost-saving exercise. It is how much science and how much national capability gets produced per unit of capital already spent. That is a different conversation from cloud cost management, and it needs different tools, a different security model, and a different definition of what a good result looks like.
About OptOps. OptOps is the optimization layer for enterprise infrastructure, built for environments where write access to production is not available. Analysis runs inside your environment, the default posture is read-only, and the same engine covers Kubernetes in the cloud and HPC and GPU compute on premise. OptOps is a DPIIT-recognised startup and a member of the NVIDIA Inception programme.
References
- Press Information Bureau, Government of India. Cabinet approves IndiaAI Mission, 7 March 2024.
- Press Information Bureau, Government of India. IndiaAI compute capacity, Lok Sabha reply, 25 March 2026.
- Press Information Bureau, Government of India. National Supercomputing Mission status, Lok Sabha reply, 10 December 2025.
- Press Information Bureau, Government of India. NSM status, Lok Sabha reply, 20 August 2025.
- Sencan, E., Kulkarni, D., Coskun, A., Konate, K. “Analyzing GPU Utilization in HPC Workloads: Insights from Large-Scale Systems.” PEARC ’25, July 2025. DOI 10.1145/3708035.3736010.

