Manager, Site Reliability Engineering

BENGALURU, IndiaFull-timePosted Jul 25, 2026

The Autonomous Recovery Service (RCV) team delivers highly available, secure, and resilient cloud services that protect Oracle Cloud Infrastructure customer data. Our mission is to provide reliable recovery capabilities that customers can depend on during planned and unplanned events.

As a hands-on technical manager, you will lead a team of engineers while also managing the reliability, operation, and continuous improvement of the RCV service. This is not a role focused solely on people management. You will remain actively engaged in production operations, technical escalations, incident response, service architecture, capacity planning, automation, and operational readiness.

You will apply your technical expertise in Oracle Database, Recovery Manager (RMAN), Zero Data Loss Recovery Appliance (ZDLRA), Exadata, and cloud infrastructure to guide technical decisions and resolve complex service issues. You will work closely with software development, database, infrastructure, and cross-functional teams to improve service reliability, scalability, security, and customer experience.

You will coach and develop engineers, establish execution priorities, review technical work, and help team members build strong operational and site reliability practices. At the same time, you will serve as a trusted escalation point for complex incidents and participate directly in troubleshooting, root cause analysis, post-incident reviews, and permanent corrective actions.

The successful candidate combines strong people leadership with deep technical judgment, operational ownership, and a willingness to stay hands-on. You should be comfortable moving between team development, customer-impacting escalations, service-level decisions, and hands-on engineering work. You will also encourage responsible use of automation and AI-assisted tools to reduce operational toil, improve productivity, and continuously evolve the RCV service.

Technical Leadership & Service Ownership

  • Lead and develop a team responsible for the reliability, operation, and continuous improvement of the Autonomous Recovery Service (RCV).
  • Remain hands-on in service operations, technical escalations, architecture reviews, incident response, and production troubleshooting.
  • Provide day-to-day technical direction, establish priorities, delegate work, and ensure team commitments are delivered.
  • Manage the end-to-end service lifecycle, including provisioning, deployments, patching, upgrades, security updates, backup and recovery, maintenance, and decommissioning.
  • Apply deep technical expertise in Oracle Database, RMAN, ZDLRA, Exadata, OCI, and related infrastructure to guide service decisions and resolve complex issues.

Capacity, Architecture & Reliability

  • Guide the team in designing and architecting reliable, secure, scalable, and highly available infrastructure and services.
  • Forecast demand, evaluate capacity requirements, and ensure RCV has sufficient resources to support current and future workloads.
  • Review service architecture, dependencies, performance characteristics, scalability, and operational readiness.
  • Partner with software development teams to deliver reliable features, infrastructure, and deployment solutions.
  • Use operational data, health reports, and performance trends to recommend improvements to service reliability, efficiency, and customer experience.

Incident Management & Escalation

  • Serve as a technical escalation point for complex incidents and issues affecting Oracle services and RCV customers.
  • Participate directly in incident response, troubleshooting, mitigation, service restoration, and cross-functional coordination.
  • Guide engineers in data collection, triage, technical analysis, debugging, and resolution of issues spanning multiple services.
  • Lead or review root cause analyses, postmortems, and corrective actions to prevent incident recurrence.
  • Ensure incidents, operational conditions, known issues, and resolutions are accurately documented.
  • Coach team members to independently perform post-incident reviews and implement permanent improvements.

Automation & Operational Excellence

  • Identify opportunities to reduce operational toil through automation, orchestration, self-service workflows, and intelligent tooling.
  • Coach team members in evaluating automation opportunities, estimating benefits, and selecting practical solutions.
  • Review automation tools, scripts, and software developed by team members for quality, safety, maintainability, and operational value.
  • Ensure automation is tested thoroughly and produces reliable, repeatable results.
  • Encourage responsible use of AI-assisted engineering and operational tools to accelerate analysis, troubleshooting, documentation, and development while maintaining sound engineering judgment.
  • Drive improvements to monitoring, alerting, observability, deployment processes, and operational readiness.

Technical Communication & Guidance

  • Enable team members to communicate the scale, capacity, security, performance, dependencies, and requirements of RCV services.
  • Review proposed infrastructure, feature, and tooling changes and help the team understand their operational and customer impact.
  • Communicate technical risks, service health, capacity needs, and improvement plans to engineering leadership and cross-functional stakeholders.
  • Build strong partnerships with software development, database, infrastructure, security, product, and business teams.
  • Promote clear documentation, effective runbooks, standard operating procedures, and knowledge sharing.

Innovation & Continuous Improvement

  • Enable engineers to experiment with new technologies and evaluate their potential impact on performance, reliability, security, and operational efficiency.
  • Supervise the execution of improvements addressing performance bottlenecks, deployment risks, resource usage, and scalability.
  • Encourage the team to challenge existing processes and recommend more effective approaches.
  • Share emerging Site Reliability Engineering, database recovery, cloud operations, automation, and AI practices with the team.
  • Use production analysis and clear data to support service design changes, investment decisions, and broader business development decisions.

Planning, Execution & Resource Management

  • Create and own execution plans for team deliverables, projects, and operational initiatives.
  • Monitor timelines, priorities, dependencies, budgets, and resource needs to ensure work is completed effectively.
  • Adjust plans as business priorities, customer needs, or operational risks change.
  • Work with leadership to identify staffing, financial, and technical resource requirements.
  • Manage operating budgets and project financials where applicable.

People Leadership & Development

  • Coach team members through technical challenges, production incidents, architectural decisions, and career development opportunities.
  • Set clear goals and expectations aligned with RCV and broader organizational objectives.
  • Identify skill gaps and provide training, mentoring, documentation, and hands-on learning opportunities.
  • Foster a culture of ownership, accountability, inclusion, continuous learning, and knowledge sharing.
  • Provide performance guidance and feedback in accordance with management processes.
  • Lead candidate interviews, contribute to talent acquisition, assess promotion readiness, and support talent planning.
  • Develop engineers who can build, test, deploy, operate, and continuously improve mission-critical cloud services.
Minimum Job Qualifications
Education and/or Experience:
8 years of experience in software engineering, infrastructure management, or related field

OR

Bachelor's Degree in Computer Science, Engineering, or related field AND 4 years of experience in software engineering, infrastructure management, or related field

OR

Master's Degree in Computer Science, Engineering, or related field AND 2 year of experience in software engineering, infrastructure management, or related field.

Job Skills:
Data Analysis Demonstrated ability to analyze and interpret data to produce actionable business insights.

Automation Experience:
3 years of experience in automation.

Programming Experience:
3 years of experience in programming and/or scripting.

Preferred Job Qualifications
Education and/or Experience:
9 years of experience in software engineering, infrastructure management, or related field

OR

Bachelor's Degree in Computer Science, Engineering, or related field AND 5 years of experience in software engineering, infrastructure management, or related field

OR

Master's Degree in Computer Science, Engineering, or related field AND 3 years of experience in software engineering, infrastructure management, or related field

OR

Doctorate in Computer Science, Engineering, or related field AND 1 year of experience in software engineering, infrastructure management, or related field.

People Leadership / Management Experience:
1 year of experience in a leadership role with or without direct reports.

Budget Experience:
1 year of experience working with operating budgets and/or project financials.

Automation Experience:
5 years of experience in automation.

Programming Experience:
5 years of experience in programming and/or scripting.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free