Senior Site Reliability Engineer

IndiaPosted Aug 5, 2026

Reliability: Ensure the reliability, scalability, and security of AI infrastructure supporting HPC & AI workloads. Incident Management: Lead incident response, root cause analysis, and continuous improvement to minimize downtime and optimize service availability. Performance Optimization: Identify and resolve bottlenecks in compute, storage, networking, and specialized hardware (GPUs, InfiniBand) to enhance AI system performance. Infrastructure Automation: Develop and maintain automation tools for deployment, monitoring, predictive analysis and management of AI infrastructure, including containerized environments (Kubernetes, Docker). Technical Leadership: Provide technical guidance in cloud and AI infrastructure technologies, collaborating with cross-functional teams to drive innovation and best practices. Customer Advocacy: Act as a customer advocate, focusing on service excellence and live site reliability for AI workloads. Research & Innovation: Stay informed on emerging AI infrastructure technologies and industry trends, recommending adoption where beneficial. Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 8+ years technical experience in software engineering, network engineering, or systems administration 12+ years of professional software engineering experience, with 8+ years in service operations, monitoring, and reliability improvement for infrastructure. 5+ years of hands-on experience developing and supporting infrastructure services for AI or cloud platforms. 1+ years experience with incident management and reliability engineering in cloud or AI environments. Doctorate Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Master's Degree in Computer Science, Information Technology, or related field AND 6+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 8+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience. 3+ years technical experience working with large-scale cloud or distributed systems. 1+ years experience in distributed systems and/or cloud platforms (Azure, Kubernetes, Docker, containers ecosystem). 1+ years experience with large supercomputers and AI platforms. 1+ years experience with GPUs, InfiniBand, or similar high-performance technologies. Publications and/or certifications related to cloud or AI infrastructure technologies a plus.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free