Senior Backend Engineer
About the Role
We're a Series A MLOps and agentic AI platform company building enterprise-grade infrastructure to help organizations deploy, manage, and monitor machine learning models at scale. Our Kubernetes-native platform powers the full ML lifecycle — from model deployment and monitoring to explainability and governance — and is now expanding into agentic AI capabilities.
We're looking for a Senior Backend Engineer with strong DevOps expertise to scale our platform and infrastructure. You'll own backend services powering agentic simulations and build tooling that helps ML researchers ship models into production. This role sits at the intersection of platform reliability, performance, and secure deployment, with close day-to-day collaboration across ML teams.
This is a hybrid role based in San Francisco, CA. Visa sponsorship is not available.
What You'll Do
Design and develop core backend services supporting distributed simulation orchestration and ML model inference.
Architect systems for task queues, distributed simulations, and scalable ML model serving.
Lead DevOps efforts, including CI/CD pipelines, infrastructure as code, observability, logging, and secure deployments.
Manage GPU workloads and model-serving pipelines to optimize throughput.
Optimize system performance for fast, concurrent simulations across agents and clusters.
Collaborate closely with ML engineers and researchers to productionize models and internal tools.
What We're Looking For
Required:
5+ years of backend development experience in production systems.
Strong DevOps experience, including designing and maintaining CI/CD pipelines and infrastructure as code (e.g., Terraform, Ansible).
Experience deploying secure, reliable, and scalable backend services.
Hands-on experience with observability and logging tooling (e.g., Prometheus/Grafana, ELK/EFK) to monitor and troubleshoot production systems.
Experience shipping ML models to production, or strong familiarity with MLOps practices to support ML researchers.
Experience with distributed systems and scalable backend architecture.
Excellent cross-functional collaboration and communication skills — you're comfortable working alongside ML researchers, data scientists, and product teams.
Nice to Have:
Background at or familiarity with cloud infrastructure providers or ML platform companies (e.g., AWS, GCP, Databricks, Anyscale, or similar).
Experience with Kubernetes-native deployments and GPU workload management.
Exposure to agentic AI or multi-agent simulation systems.
Compensation & Benefits
Base salary: $180,000 – $200,000 USD annually
Equity participation (standard for stage and role)
Comprehensive health, dental, and vision benefits
Collaborative, research-oriented engineering culture
Location
Primary location: San Francisco, CA, United States
Work arrangement: Hybrid (on-site + remote flexibility)
Visa sponsorship: Not available