Senior Technical Program Manager, Cluster Operations & Quota Management

United StatesPosted Aug 6, 2026

As the TPM for Cluster Operations, you own the end-to-end system that gets quota and cluster changes implemented across MAI's GPU fleet - not just the execution of any single change, but the machinery that makes every change fast, clean, and repeatable. When an allocation is decided, you own getting it implemented: preparing and sequencing the changes ahead of time, running cluster moves and cycle cutovers cleanly, coordinating with infrastructure, platform, HPC, and vendor teams to provision capacity, and validating that quota is genuinely usable rather than merely configured. You chase every dependency down, surface and clear blockers, and keep quota flowing to where it matters most. Owning the system means continuously rebuilding it. You champion the tooling squads and leadership use to see quota, utilization, and idle capacity - partnering with the engineering teams that build it so it reflects reality and so new clusters and tenants onboard smoothly as the fleet grows. And you work proactively: spotting cross-functional dependencies before they bite, streamlining the handoffs between teams, and convening the right working groups across research, infrastructure, platform, and vendor organizations to fix the operating model at its root, not just the symptom in front of you. Every cycle should be faster, cleaner, and more automated than the last, and you are the person accountable for making that true. Own the end-to-end operating system for quota and cluster execution: the process, playbooks, tooling agenda, and cross-company coordination that turn allocation decisions into usable capacity Execute approved quota and capacity allocations across MAI's squads end to end: prepare and sequence changes in advance, run cluster moves and cycle cutovers cleanly, and drive each one to completion across every partner team. Own the full chain to usable capacity, not just configured quota: coordinate provisioning and environment readiness with infrastructure, platform, HPC, and vendor teams, and validate identity, access, and utilization so researchers are productive from day one. Relentlessly chase dependencies, surface blockers early, and unblock them, keeping a crisp, auditable status on every change so nothing is dropped or misreported. We know that strong candidates come from many different backgrounds and experiences. If you are excited about the role and believe you could make an impact, we encourage you to apply, even if you do not meet every qualification listed below. Significant experience in technical program management within infrastructure, platform engineering, or other compute-intensive environments. A record of proactively improving operating models: spotting process failures, convening the right people across organizations, and driving automation and streamlining without waiting to be asked. High agency and a service-oriented mindset, comfortable relentlessly chasing and pushing across teams, without formal authority, to get things done fast.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free