MIG packing & fractionalisation
Slice and pack GPUs so multiple tenants and jobs share one card safely — more workloads per accelerator.
For HPC & AI Data Centers
OptOps reads real GPU telemetry, packs MIG slices, tunes power and clocks, and reclaims idle silicon — up to 30–38% more efficiency from the fleet you already own, without buying a single new accelerator.
6 GPUs · 12 MIG slices scattered · avg 31% SM utilisation
01 · The problem
Compute demand is exploding and power is the new ceiling. You can’t buy your way out — accelerators are scarce and the grid is capped. The only fast lever is getting more from the fleet you already own.
30–50%
typical GPU utilisation in AI clusters
— NVIDIA, 2026
$25–40K
per H100 — idle capacity burns CapEx
— Introl
#1
power is becoming the binding constraint on data-centre buildout
— Goldman Sachs
~2×
data-centre electricity consumption by 2030 — 485 → ~950 TWh
— IEA
scheduler view · card handed out whole · “busy”
silicon view · model parked in VRAM · barely computing
A GPU can hold a model in VRAM while doing almost no computing — allocation says busy, the silicon is idle. That gap is $25–40K per H100 per year (Introl, 2025).
Many scheduling setups allocate an entire GPU per job — leaving unused compute and memory capacity stranded.
Tenants reserve for peak demand and sit idle off-peak. The buffer becomes permanent waste.
A GPU can hold a model in VRAM while doing almost no computing. Allocation metrics say busy — the silicon is idle.
02 · Why today’s tools fall short
Your stack already measures utilisation and places jobs — and idle stays high. Seeing waste and safely removing it are two different jobs.
What you already have
They monitor, allocate and orchestrate — but the optimization loop is still missing.
What’s still missing
03 · Accelerator optimization
Every lever is risk-scored, opt-in, and applied under your policies and SLAs — with zero risk to running jobs.
Slice and pack GPUs so multiple tenants and jobs share one card safely — more workloads per accelerator.
Match VRAM and SM allocation to what jobs actually use — not the whole 80 GB card.
Lift actual compute activity, not just allocation — turn “loaded” GPUs into working ones.
Tune GPU power caps and clocks to each workload — large energy savings with no loss on the jobs that matter.
Place and pack jobs to run inside your power and thermal budget — never blow the cap.
Pause lower-priority jobs to free GPUs for paying demand, then resume exactly where they stopped — no recomputation, no lost hours.
04 · Power & efficiency
For a power-bound data center, efficiency is capacity — every watt saved is a watt you can sell.
Serving compute
420 kW
Idle + untuned burn
310 kW
Cooling overhead
210 kW
Sellable headroom
60 kW
Illustrative split for a 1 MW hall — actual gains depend on workload mix and are measured per deployment.
30–38%
less energy
Measured in deployment — from power/clock tuning + MIG packing.
~41%
of AI-cluster power goes to accelerators
And they keep drawing it even when idle (research, 2026) — tuning GPU power is the single biggest efficiency lever you have.
05 · How it works
Sits alongside your stack, on-prem or air-gapped. Automation is opt-in, under your policies.
Air-gap ready
Read-only agent, air-gapped deployment. Sits alongside your HPC stack — no infrastructure changes to start.
First scan
Reads DCGM and accelerator telemetry — real SM and memory utilisation and power draw, per GPU and per job.
Ongoing
MIG packing, power and clock tuning, and right-sizing recommendations — each tagged with ROI and a confidence score.
First week
Apply safely under your policies and SLAs. Up to 30–38% more efficiency from the fleet you already own.
write access exists only for approved changes — scoped per action, temporary, revocable, fully logged
We started where the world’s GPUs actually run — the scheduler behind virtually every supercomputer and national research cluster.
The same engine, extending to K8s GPU fleets — backed by our production-proven Kubernetes platform. Design partners get first access.
06 · Safe & sovereign
The first question every operator asks: “what if this touches a tenant’s job — or their data?” The architecture is the answer.
The agent needs no write access to deliver value — monitoring reads metrics, specs and usage only.
Each recommendation carries a risk score; nothing low-confidence is ever proposed for automation.
You approve every change. Write access is scoped per action, temporary, revocable and fully logged.
Any applied change can be reverted instantly — production safety is the architecture, not a feature.
Runs fully inside sovereign, air-gapped environments — built for national labs and residency-bound compute.
Only infrastructure telemetry is read — never tenant applications, models, or data.
07 · Proof
30–38%
energy saved
power/clock tuning + MIG packing, measured in deployment
MIG + spot
demonstrated live
autonomous right-sizing of jobs on real GPUs
National HPC
supercomputing deployment
accelerator/HPC proof with an authorised national supercomputing programme
Where we actually are — the honest version
Live and proven on Kubernetes + GPU, delivering real accelerator utilisation and power savings today. Deployed with an authorised national supercomputing programme as our accelerator/HPC proof. Preemption and checkpointing are in design-partner beta — we’d prove them on your cluster together, on your workloads, before you commit. We don’t overclaim a finished HPC product: we bring a proven engine and prove it on your fleet.
08 · The economics
For a data center that sells capacity, every point of utilisation is revenue. Sweat the asset.
Pack + right-size → more tenants per GPU → more revenue from the same fleet. No new hardware.
Efficiency frees watts → deploy and sell more compute inside the same envelope. For a power-bound facility, efficiency is capacity.
Get more from installed accelerators before the next multi-million-dollar cluster — better ROIC per rack.
handshakeCommercial model: pay on verified savings and gains. Our upside is tied to yours — no savings, no bill.
09 · Roadmap
One engine, expanding — Kubernetes to GPU to HPC to every substrate that runs AI.
10 · Next steps
Preemption + checkpointing + power optimization on a scoped slice of your fleet — measure the real gain before you commit.
An infrastructure analysis that identifies high-impact optimization opportunities and quantifies potential savings in dollars and watts.
Power-consumption and accelerator-optimization detail, and a pay-on-verified-gains model agreed in writing before anything begins.
We’ll walk your cluster view, resource footprint, and cost insights live — then scope a beta on a slice of your fleet and measure the real gain.
or contact us — founder-led, operator-built
Running Kubernetes in the cloud too? The same engine does node consolidation, spot automation, and live migration.