Reliability: Ensure the reliability, scalability, and security of AI infrastructure supporting HPC & AI workloads. Incident Management: Lead incident response, root cause analysis, and continuous improvement to minimize downtime and optimize service availability. Performance Optimization: Identify and resolve bottlenecks in compute, storage, networking, and specialized hardware (GPUs, InfiniBand) to enhance AI system performance. Infrastructure Automation: Develop and maintain automation tools for deployment, monitoring, predictive analysis and management of AI infrastructure, including containerized environments (Kubernetes, Docker). Technical Leadership: Provide technical guidance in cloud and AI infrastructure technologies, collaborating with cross-functional teams to drive innovation and best practices. Master's Degree in Computer Science, Information Technology, or related field AND 1+ year(s) technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 6+ years technical experience in software engineering, network engineering, or systems administration 8+ years of professional software engineering experience, with 5+ years in service operations, monitoring, and reliability improvement for infrastructure. 1+ years experience with incident management and reliability engineering in cloud or AI environments. Master's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 5+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience. 2+ years technical experience working with large-scale cloud or distributed systems. 1+ years experience in distributed systems and/or cloud platforms (Azure, Kubernetes, Docker, containers ecosystem). 1+ years experience with GPUs, InfiniBand, or similar high-performance technologies.
Want jobs like this matched to you?
SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.