Senior System Software Engineer - Dynamo-Triton Inference Server
Role responsibilities
Develop GPU-accelerated AI inference serving software and drive the convergence of Triton Inference Server and NVIDIA Dynamo stacks. Balance production robustness with the optimization of prediction throughput and latency for LLM and non-LLM workloads.
Requirements
Requires a Master's or PhD in Computer Science or equivalent experience with over 5 years of professional deep learning software experience. Proficiency in Rust, C++, and Python is essential, along with experience in high-scale distributed ML systems.
Key skills
Rust, C++, Python, Deep Learning, Distributed Systems, Software Design, Performance Analysis, Debugging, Test Design, GPU Memory Management, Cache Management, High-performance Networking, TensorRT, PyTorch, ONNX, vLLM
Keywords
GPU, AI Inference, Triton Inference Server, NVIDIA Dynamo, Large Language Models, LLM, Open Source, TensorRT, PyTorch, ONNX, OpenVINO, vLLM, TRT-LLM, Distributed Systems, Rust, C++, Python, GitHub, OSS Licensing, Deep Learning, Software Engineering, Performance Tuning, Cloud Environments