Research Member of Technical Staff - Training Platform

Mountain View, CAPosted May 19, 2026
Research Member of Technical Staff - Training Platform LocationMountain ViewEmployment TypeFull timeDepartmentResearchCompensation$200K – $300K • Offers EquityAt Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.We're looking for a Research Engineer to build and maintain the training platform that powers our model development — experiment orchestration, job management, observability, and the tooling that lets researchers move from idea to result as fast as possible.What You'll DoBuild and maintain training orchestration systems for large-scale distributed model training across GPU clustersDevelop experiment management tooling: job configuration, tracking, reproducibility, and artifact managementBuild observability infrastructure for training runs: loss curves, compute utilization, gradient statistics, and anomaly detectionOptimize and automate the research iteration loop from experiment launch to results analysisManage job scheduling and cluster utilization for efficient use of GPU computeBuild internal tooling and interfaces that help researchers move fasterCollaborate with training systems, data infrastructure, and research teams to support their platform needsWhat We're Looking ForStrong software engineering skills with experience in MLOps or ML platform engineeringFamiliarity with distributed training frameworks (PyTorch DDP, FSDP, DeepSpeed, Megatron, or similar)Experience building experiment tracking, reproducibility, and artifact management systemsComfortable managing and operating GPU cluster environments (Slurm, Kubernetes, or similar)Strong reliability engineering instincts: monitoring, alerting, and failure recoveryNice to Have (But Not Required)Experience with training orchestration tools (Slurm, Ray, Kubernetes, or similar schedulers)Familiarity with experiment tracking tools (Weights & Biases, MLflow, or custom solutions)Experience supporting large model training pipelines (LLMs, VLMs, or video models)Understanding of parallelism strategies and how they affect training efficiency and debuggingExperience with cloud-based training infrastructure (AWS, GCP, or Azure)Why This RoleYour platform is the daily tool every researcher and engineer uses to train modelsImprovements to training velocity and reliability compound across every experiment the team runsHigh visibility with direct feedback from researchers and ML engineersBuild systems that scale from today's models to future frontier training runsApply for this Job

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free