Research Member of Technical Staff- Efficient Modeling
Mountain View, CAPosted May 19, 2026
Research Member of Technical Staff- Efficient Modeling LocationMountain ViewEmployment TypeFull timeDepartmentResearchAt Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.We're looking for a Research Scientist or Research Engineer focused on model efficiency — making our foundation world models faster, smaller, and more deployable without sacrificing capability. This work is critical to closing the gap between research-scale models and real-time operation on robot hardware.What You'll DoResearch and implement model compression techniques: quantization, pruning, structured sparsity, distillation, and low-rank approximationDesign efficient architectures and attention mechanisms suited to real-time inference on edge and robot hardwareDevelop training strategies that produce better accuracy-efficiency tradeoffs from the startProfile and benchmark models across hardware targets to identify and resolve efficiency bottlenecksBuild evaluation frameworks that measure capability retention after compression or architecture changesCollaborate with training systems and deployment teams to ensure efficient models translate to faster real-world inferencePublish and present work at top-tier venues What We're Looking ForStrong understanding of model compression and efficient architectures for large modelsHands-on experience with quantization, distillation, or pruning applied to transformers or large neural networksDeep knowledge of where efficiency gains are possible in modern architecturesProficiency with PyTorch and familiarity with hardware-aware optimization (CUDA, TensorRT, or similar)Ability to run principled experiments that characterize capability-efficiency tradeoffsNice to Have (But Not Required)PhD in ML, CS, or a related field — or equivalent research/engineering experiencePublication record at NeurIPS, ICML, ICLR, MLSys, or related venuesExperience with efficient video or multimodal model architecturesFamiliarity with edge deployment targets (Jetson, custom ASICs, or mobile hardware)Prior work on speculative decoding, early exit, or adaptive computeExperience deploying compressed models on physical robots or latency-constrained systemsWhy This RoleBridge the gap between large-scale research models and real-time robot deploymentsYour work determines whether frontier capabilities actually run on our hardwareHigh leverage: efficiency improvements benefit every model the team trains and deploysWork at a rare intersection of deep learning research and systemsApply for this Job