Senior Site Reliability Engineer

BENGALURU, IndiaFull-timePosted Jul 25, 2026

The Autonomous Recovery Service (RCV) team is responsible for delivering highly available, secure, and resilient cloud services that protect Oracle Cloud Infrastructure (OCI) customer data. Our mission is to ensure customers can confidently recover from planned and unplanned events through industry-leading recovery capabilities, intelligent automation, and operational excellence.

As a Software Development & Site Reliability Engineer, you will play a critical role in operating and continuously improving mission-critical cloud services. This is an operations-first engineering role where you'll own the health, reliability, and lifecycle of production services while developing software and automation that reduce operational complexity, improve service resilience, and enhance the customer experience.

You'll work across the full service lifecycle—from deployment, patching, upgrades, monitoring, incident response, and root cause analysis to designing automation, improving observability, and implementing engineering solutions that eliminate repetitive operational work. Success in this role requires curiosity, strong analytical thinking, and a passion for solving complex operational challenges through software and automation.

Our engineers embrace AI as a force multiplier, using AI-assisted development and operational tools to accelerate problem solving, improve productivity, and build smarter, more autonomous systems. We value engineers who combine technical expertise with critical thinking, sound engineering judgment, and a continuous improvement mindset to challenge existing processes and drive innovation.

If you enjoy owning production services, building reliable cloud infrastructure, automating everything possible, and working on technology that protects mission-critical customer data at cloud scale, we'd love to have you join our team.

  • Support the day-to-day operation, health, availability, and reliability of Oracle's Autonomous Recovery Service (RCV) production environments.
  • Perform operational activities including deployments, patching, upgrades, security updates, infrastructure maintenance, and production change management following established operational procedures.
  • Monitor production services, investigate alerts, troubleshoot issues, and assist in restoring service while helping the team achieve Service Level Objectives (SLOs).
  • Participate in an on-call rotation with guidance from senior engineers to support production services and ensure a high level of customer availability.
  • Assist in incident response, root cause analysis (RCA), and post-incident reviews, implementing corrective actions to prevent recurring issues.
  • Develop, test, and maintain automation tools, scripts, and software that simplify operational tasks, improve reliability, and reduce operational toil.
  • Contribute to Infrastructure as Code (IaC) solutions using Terraform to automate provisioning, configuration, and management of Oracle Cloud Infrastructure (OCI) resources.
  • Collaborate with Software Development and Site Reliability Engineering teams to build reliable, scalable, and maintainable cloud services.
  • Analyze logs, metrics, and operational data to identify trends, troubleshoot issues, and recommend improvements to service performance and reliability.
  • Support capacity planning, performance optimization, and service scalability initiatives.
  • Develop and enhance monitoring, alerting, dashboards, and observability solutions that improve operational visibility.
  • Assist in managing Oracle Database environments, including backup, restore, recovery validation, and disaster recovery operations using Oracle Recovery Manager (RMAN).
  • Write and maintain automation using Python, Bash, REST APIs, and other scripting technologies to improve operational efficiency.
  • Use Git and Bitbucket to manage source code, collaborate on development efforts, and participate in modern CI/CD workflows.
  • Leverage AI-assisted engineering tools to accelerate software development, operational analysis, troubleshooting, and documentation while applying sound engineering judgment to validate results.
  • Identify opportunities to automate repetitive operational processes and implement intelligent solutions that improve productivity, service reliability, and customer experience.
  • Document operational procedures, troubleshooting guides, and lessons learned to improve team knowledge and operational readiness.
  • Collaborate effectively with software engineers, cloud infrastructure teams, database engineers, and cross-functional partners to deliver reliable cloud services.
  • Continuously expand technical knowledge by learning Oracle Cloud Infrastructure (OCI), Oracle Database, RCV architecture, Site Reliability Engineering practices, Infrastructure as Code (Terraform), automation, and AI-assisted engineering techniques.
  • Demonstrate curiosity, critical thinking, and a continuous improvement mindset by challenging existing processes and contributing innovative ideas that improve service reliability, operational excellence, and customer outcomes.

Technologies You'll Work With

  • Cloud & Infrastructure: Oracle Cloud Infrastructure (OCI), Linux, Terraform (Infrastructure as Code)
  • Database Technologies: Oracle Database, Oracle Recovery Manager (RMAN), backup and recovery, disaster recovery
  • Development & Automation: Python, Bash/Shell scripting, REST APIs, JSON, YAML
  • Source Control & DevOps: Git, Bitbucket, CI/CD pipelines
  • Site Reliability Engineering: Production monitoring, observability, incident response, root cause analysis, performance optimization, capacity planning
  • Modern Engineering: AI-assisted development, intelligent automation, operational analytics, Agile engineering practices

 

Minimum Job Qualifications
Education and/or Experience:
8 years of experience in software engineering, infrastructure management, or related field

OR

Bachelor's Degree in Computer Science, Engineering, or related field AND 4 years of experience in software engineering, infrastructure management, or related field

OR

Master's Degree in Computer Science, Engineering, or related field AND 2 year of experience in software engineering, infrastructure management, or related field.

OR

Doctorate in Computer Science, Engineering, or related field

Job Skills:
Same skills as prior level plus;
Operating Systems Demonstrated ability in or knowledge of operating systems, including installing, upgrading, and troubleshooting various operating environments.

Automation Experience:
3 years of experience in automation.

Programming Experience:
3 years of experience in programming and/or scripting.

Preferred Job Qualifications
Education and/or Experience:
9 years of experience in software engineering, infrastructure management, or related field

OR

Bachelor's Degree in Computer Science, Engineering, or related field AND 5 years of experience in software engineering, infrastructure management, or related field

OR

Master's Degree in Computer Science, Engineering, or related field AND 3 years of experience in software engineering, infrastructure management, or related field

OR

Doctorate in Computer Science, Engineering, or related field AND 1 year of experience in software engineering, infrastructure management, or related field.
Automation Experience:
5 years of experience in automation.
Programming Experience:
5 years of experience in programming and/or scripting.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free