Research Scientist, Performance Engineering

San Francisco, CAPosted Jul 13, 2026
Research Scientist, Performance Engineering LocationSan FranciscoEmployment TypeFull timeLocation TypeOn-siteDepartmentEngineeringCompensation$200K – $300K • Offers EquityTBC is building next-generation AI systems at the intersection of biological computing, generative models, and large-scale AI infrastructure. As we scale our world-model and neural-optimizer efforts, we are looking for an optimization-focused Research Scientist / ML Engineer to improve the efficiency, latency, throughput, and deployability of large models.This role is focused on making frontier models run faster, cheaper, and more reliably — especially LLMs, diffusion models, video generation models, and world-model systems. You will work across inference optimization, training efficiency, model compression, memory management, and GPU-level performance to help turn research systems into scalable, customer-ready products.What You’ll Work OnOptimize inference for LLMs, diffusion models, video models, and world-model systemsImprove serving efficiency through techniques such as KV caching, batching, quantization, distillation, speculative decoding, and memory optimizationBuild and optimize high-throughput inference pipelines for large models running on GPU clustersProfile model performance across latency, throughput, memory usage, GPU utilization, and costImplement custom kernels or low-level optimizations using Triton, CUDA, PyTorch, or related systemsImprove training and fine-tuning efficiency for large generative models, including distributed training, checkpointing, parallelism, and data loadingWork with research teams to identify bottlenecks in model architecture, inference paths, and deployment workflowsTranslate model performance improvements into clear customer-facing benchmarks and technical proof pointsEvaluate trade-offs across model quality, latency, cost, memory, and deployabilityWhat We’re Looking ForStrong background in machine learning systems, model optimization, or high-performance AI infrastructureHands-on experience optimizing LLMs, diffusion models, video generation models, or other large generative systemsExperience with one or more of:Inference optimizationKV caching / attention optimizationTriton or CUDA kernel developmentQuantization, pruning, distillation, or model compressionDistributed training / fine-tuning efficiencyGPU profiling and performance debuggingStrong PyTorch experience and comfort working close to the model/runtime boundaryAbility to reason about trade-offs between quality, latency, throughput, memory, and costComfortable working across research code, production systems, and benchmarking infrastructureExcited to work in an ambiguous, early-stage environment where optimization work directly shapes product feasibilityWhat Success Looks LikeLarge models run faster, cheaper, and more reliably across TBC’s core workloadsInference pipelines show measurable improvements in latency, throughput, memory use, and GPU utilizationTraining and fine-tuning workflows become more efficient, reproducible, and scalableOptimization work translates into clear product and customer value, not just internal benchmarksResearch prototypes become deployable systems that can support demos, evaluations, and early partner use casesPreferred QualificationsPhD, MS, or equivalent industry experience in Computer Science, Machine Learning, Systems, Robotics, or related fieldPrior work optimizing large-scale generative models in production or research settingsExperience with modern inference/training stacks such as PyTorch, Triton, CUDA, vLLM, TensorRT, DeepSpeed, FSDP, Ray, or similar toolingExperience working with LLMs, diffusion models, video generation models, or world modelsApply for this Job

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free