Software Engineer - Compute Infra / HPC

United StatesPosted Aug 5, 2026

Design and build distributed services, control planes, and platform APIs for the end-to-end lifecycle of AI compute clusters, from capacity ingestion and provisioning through certification, upgrades, maintenance, repair, and decommissioning. Build foundational primitives that make new capacity enablement repeatable and secure by default, including cluster bootstrap, Kubernetes control planes, networking, identity and access, image distribution, secrets, infrastructure as code, and cluster metadata across Azure and partner clouds. Own software systems that reduce cluster bring-up time by automating rack and node qualification, topology validation, scale testing, and the transition from newly delivered hardware to healthy, schedulable fleet. Develop the node and fleet health platform that combines telemetry, health signals, diagnostics, and lifecycle state to detect, isolate, and remediate unhealthy hardware, hosts, networks, and storage with minimal human intervention. Build policy-driven and agent-assisted automation for certification, maintenance, safe rollout or rollback of software and platform components. Use production evidence, benchmarks, workload profiles, and fault injection to improve system architecture, efficiency, failure detection, recovery time, and the quality of health and readiness decisions. Translate incidents and repeated operational work into durable software, stronger abstractions, automated guardrails, and clear ownership boundaries rather than one-off manual procedures. Provide technical leadership through design reviews, architecture decisions, code quality, mentoring, and execution of complex cross-team initiatives. Embody our Culture and Values. Bachelor's Degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field AND 4+ years of software engineering experience building production distributed systems, infrastructure platforms, or control-plane services 4+ years of experience designing and implementing scalable software for cloud, datacenter, cluster, or fleet infrastructure using one or more general-purpose languages such as Go, Rust, C++, C#, Java, or Python. Experience with Kubernetes or comparable cluster-orchestration systems, Linux systems, public-cloud infrastructure, and infrastructure-as-code or declarative configuration. Demonstrated ability to debug complex behavior across service, control-plane, operating-system, networking, and hardware boundaries and turn findings into robust software improvements. Experience leading projects across multiple teams and communicating technical tradeoffs clearly to engineering and non-engineering stakeholders. Master's Degree in Computer Science or a related technical field AND 6+ years of software engineering experience building distributed systems or infrastructure platforms OR equivalent experience. Experience building or managing compute infrastructure at hyperscale, such as fleets with thousands of nodes, multiple clusters, heterogeneous accelerator generations, or capacity spanning multiple cloud and datacenter providers. Experience with cluster and node lifecycle systems, including bare-metal provisioning, cluster bring-up, health checking, certification, maintenance, repair automation, capacity management, or decommissioning. Depth in Kubernetes internals or related orchestration technologies, including controllers and operators, scheduler, autoscaler, kubelet, Cluster API, container runtimes, or workflow engines. Experience building cloud foundations, including virtual networking, CNI, private connectivity, DNS, workload identity, RBAC, secrets, image and artifact distribution, and Terraform or equivalent infrastructure-as-code systems. Experience designing health, diagnostics, and remediation platforms using event-driven architectures, durable workflows, state machines, telemetry pipelines, automated decisioning, and safe recovery mechanisms. Low-level systems experience with Linux kernels, device drivers, firmware, BMC or Redfish, secure boot or attestation, virtualization, or hardware-health agents. Experience with GPU or accelerator systems and high-performance interconnects such as InfiniBand, RoCE or RDMA, NVLink, NCCL or collective communication, as well as high-performance storage and data movement. Experience owning production reliability and incident response while systematically eliminating recurring toil through software and automation. Dedication to writing clean, testable, maintainable, secure, and well-documented code, with strong engineering judgment around correctness, scalability, operability, and long-term platform evolution. Ability to create alignment across senior stakeholders, mentor engineers, and drive complex multi-quarter technical initiatives from architecture through production adoption. Passion for learning rapidly evolving infrastructure technologies and for making enormous compute systems more reliable, efficient, understandable, and easy to use.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free