Software Engineer, Inference Runtime

EngRadar
New York, NYFULL_TIMEPosted Aug 9, 2026

Role responsibilities

The role involves developing and optimizing the inference stack for on-device and cloud AI, integrating new engines and multimodal models. Responsibilities include improving latency and throughput across various hardware runtimes and contributing to open-source projects.

Requirements

Candidates need significant experience building production ML systems or performance-sensitive infrastructure with strong proficiency in Python and C++. Deep knowledge of transformer architectures and experience profiling CPU/GPU workloads are essential.

Key skills

Python, C++, Transformer Architectures, Model Inference, CPU Profiling, GPU Profiling, PyTorch, Llama.cpp, MLX, ExecuTorch, vLLM, SGLang, TensorRT-LLM, CUDA, Metal, Vulkan

Keywords

Inference Runtime, Large Language Models, LLM, On-device AI, Transformer Architectures, PyTorch, Llama.cpp, MLX, ExecuTorch, vLLM, SGLang, TensorRT-LLM, CUDA, Metal, Vulkan, ROCm, CPU, GPU, Model Loading, Batching, Scheduling, Caching, Distributed Execution, Performance Profiling, Open-source, Multimodal Models

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free