Senior System Software Engineer - Dynamo-Triton Inference Server

Thomas To
Santa Clara, CAFULL_TIMEPosted Aug 6, 2026

Role responsibilities

Develop GPU-accelerated AI inference serving software and drive the convergence of Triton Inference Server and NVIDIA Dynamo stacks. Balance production robustness with the optimization of prediction throughput and latency for LLM and non-LLM workloads.

Requirements

Requires a Master's or PhD in Computer Science or equivalent experience with over 5 years of professional deep learning software experience. Proficiency in Rust, C++, and Python is essential, along with experience in high-scale distributed ML systems.

Key skills

Rust, C++, Python, Deep Learning, Distributed Systems, Software Design, Performance Analysis, Debugging, Test Design, GPU Memory Management, Cache Management, High-performance Networking, TensorRT, PyTorch, ONNX, vLLM

Keywords

GPU, AI Inference, Triton Inference Server, NVIDIA Dynamo, Large Language Models, LLM, Open Source, TensorRT, PyTorch, ONNX, OpenVINO, vLLM, TRT-LLM, Distributed Systems, Rust, C++, Python, GitHub, OSS Licensing, Deep Learning, Software Engineering, Performance Tuning, Cloud Environments

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free