##
Company:
Qualcomm India Private Limited
## Job Area:
Engineering Group, Engineering Group > Software Engineering
General Summary:
We are looking for a Senior Site Reliability Engineer to join a 24/7, follow-the-sun SRE team responsible for keeping our critical production systems reliable, scalable, and secure around the clock. As part of a globally distributed team spanning multiple time zones, you will share on-call coverage that follows daylight hours rather than night shifts, and partner closely with development teams to automate operations, strengthen observability, and continuously improve system resilience.
Main Responsibilities:
* System Reliability: Ensure the reliability, availability, and performance of critical systems.
* Service Level Objectives: Define, measure, and report on SLIs, SLOs, and error budgets, and use them to prioritize reliability work.
* Automation: Develop and maintain automation scripts and tools to streamline operations.
* Monitoring: Develop and maintain monitoring dashboards & alerts
* Incident Management: Lead incident response efforts and post-mortem analysis to prevent future occurrences.
* Performance Tuning: Optimize system performance and scalability.
* Cost Optimization: Monitor and optimize cloud spend (e.g., right-sizing and autoscaling with Karpenter) to balance reliability with cost efficiency.
* Security: Implement and maintain security best practices.
* Documentation: Create and maintain comprehensive documentation for systems and processes.
* Mentorship: Mentor junior engineers and champion SRE best practices across cross-functional teams.
* On-Call Shifts: Own front-line 24/7 on-call rotations and incident response for critical production systems, acting as the reliability shield for the platform and observability engineering teams so they are not paged for production incidents.
Requirements:
Qualification:
* Education: Bachelor’s degree in computer science, Engineering, or a related field. Advanced degrees are a plus.
* Experience: 5+ years of experience in a similar role, with a strong background in software engineering and systems administration.
Technical Skills:
* Programming Languages: Proficiency in one or more programming languages such as Python, Go.
* Cloud Platforms: Extensive experience with AWS cloud platform.
* Infrastructure as Code: Hands-on experience with tools like Terraform, Ansible, or CloudFormation.
* Containerization and Orchestration: Expertise in Docker and Kubernetes. Kubernetes - Hands-on experience is a must.
* Kubernetes technologies (with preferences): ArgoCD, Linkerd, Prometheus, Karpenter, etc…
* Monitoring and Logging: Proficiency with monitoring tools like Prometheus, Grafana, and logging tools like ELK stack or Loki stack.
* CI/CD Pipelines: Experience with continuous integration and continuous deployment tools such as Jenkins, GitLab CI, or CircleCI.
* Networking: Strong understanding of networking concepts, protocols, and security.
Minimum Qualifications:
• Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.
OR
Master's degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience.
OR
PhD in Engineering, Information Systems, Computer Science, or related field.
• 2+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc.
Soft Skills:
* Problem-Solving: Excellent analytical and troubleshooting skills.
* Communication: Strong verbal and written communication skills.
* Collaboration: Ability to work effectively in a team environment and collaborate with cross-functional teams.
* Leadership: Proven leadership skills and the ability to mentor junior engineers.
* Remote Work: Comfortable working in a fully distributed, offshore setup and collaborating effectively with...
Want jobs like this matched to you?
SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.