Senior Software Engineer, ML Infrastructure
Cupertino$209k–$235kPosted Aug 3, 2026
The Role
We're looking for a Senior Software Engineer to join our ML Infrastructure team and support the foundational infrastructure that powers Gridmatic. Our platform challenges are shaped by the nature of energy markets: forecasts and trading decisions run on tight schedules, battery dispatch commands must execute reliably in real time, and ML models need to train and deploy continuously as new data arrives.
What You'll Do
- Design and build the compute platform that runs Gridmatic's production services, batch jobs, and ML training workloads
- Help own our production workflow orchestration end to end, from the cluster and node pools it runs on to the abstractions our teams build on top of it
- Make our ML iteration cycle faster and cheaper by profiling workflows, cutting latency, and improving how we use compute
- Build the observability and developer tooling that lets engineers understand and troubleshoot their own workloads, from post-run cost reporting to monitoring and alerting
- Improve the reliability and resiliency of production workloads, including how we handle capacity constraints and multi-region routing
- Manage GPU and accelerator capacity: node pools, drivers, spot vs. on-demand tradeoffs, and scheduling and quota so training and batch jobs get the compute they need without overspending
- Own autoscaling and quota management so the platform scales up under load and down to zero when idle
- Work closely with the ML team to find pain points and quickly ship solutions
- Make architectural decisions that shape how we build software as we grow
What We're Looking For
- Significant experience building and operating production infrastructure on a public cloud platform (we run on GCP, but AWS or Azure experience translates well)
- Strong distributed systems and infrastructure skills: standing up services, scaling and debugging Kubernetes (we run on GKE and it's foundational to our platform), writing Terraform, and comfort with cloud networking, IAM, and secrets management
- Hands-on experience with workflow orchestration tools like Flyte, Temporal, or Airflow
- Proficiency in Python, and either already know Go or have experience with a similar systems language (C++, Java, Rust) and are excited to work in Python and Go day-to-day
- Experience working closely with ML engineers or data scientists as your customers, and genuine excitement to keep doing it
- Clear communication, whether writing a design doc, reviewing code, or explaining a complex system to someone new to it
Nice to Have
- Experience running GPU or accelerator workloads on Kubernetes (node pools, drivers, scheduling, quota) is strongly preferred for this role
- Familiarity with observability tooling (we use Grafana + Google Cloud Monitoring)
- Experience building internal developer platforms or tooling that other engineers rely on
- Prior work in domains where latency and reliability have direct business consequences