Design, implement, and operate cloud-scale services that improve GPU inference efficiency, utilization, and cost across Azure. Partner with Azure Systems Research and Azure AI infrastructure teams to translate algorithms, models, and systems research into production-ready infrastructure. Build telemetry, experimentation, and control-plane capabilities that identify inference-efficiency opportunities and safely validate improvements at fleet scale. Write high-quality production code in languages such as C#, Rust, Python, or C++ with engineering practices for reliability, security, and maintainability. Drive architecture and technical strategy for workload placement, capacity management, performance isolation, and optimization of inference-serving systems. Collaborate across Azure Core, hardware, platform, and artificial intelligence teams to land cross-organizational initiatives in production. Mentor engineers, lead technical reviews, and help build an inclusive engineering culture focused on customer impact and operational excellence. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. 4+ years of cloud infrastructure or distributed systems development experience. 2+ years of experience developing, deploying, or optimizing machine learning, artificial intelligence, or inference-serving systems at scale. 4+ years of experience analyzing and troubleshooting large-scale distributed systems in critical production online service environments, including experience with GPU infrastructure, machine learning inference systems, workload scheduling, capacity optimization, or performance engineering.
Want jobs like this matched to you?
SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.