Vice President, Enterprise Technology Command Center, Major Incident Manager, DTI Site Reliability Engineering
Business Function
Group Technology enables and empowers the bank with an efficient, nimble and resilient infrastructure through a strategic focus on productivity, quality & control, technology, people capability and innovation. In Group Technology, we manage the majority of the Bank's operational processes and inspire to delight our business partners through our multiple banking delivery channels.
Role Purpose:
VP, Major Incident Manager for ETCC SRE Operations, is responsible for the end-to-end management of critical and major incidents impacting DBS's technology services. This role focuses on minimizing service disruption, restoring service rapidly, and driving continuous improvement in incident response processes. The Major Incident Manager will lead cross-functional teams during incidents, ensuring clear communication, effective problem resolution, and adherence to SRE principles for reliability and operational excellence.
Key Responsibilities:
- Major Incident Management:
- Lead and manage major incidents from detection through resolution, ensuring rapid service restoration and minimal business impact.
- Act as the primary communication point during major incidents, providing timely and accurate updates to senior management, stakeholders, and affected business units. Own incident escalation to senior management
- Coordinate and drive technical teams (including SRE, development, infrastructure, and operations) to diagnose, troubleshoot, and resolve complex production issues.
- Ensure all incidents are thoroughly documented, including timelines, actions taken, and resolution steps.
- Ensure compliance with MAS, local in-country regulatory process during incident reporting requirements
- Facilitate Blameless-incident reviews (BIRs) to identify root causes, preventive measures, and opportunities for process and system improvements.
- Experience in generating management reports on incidents, problem trends, and thematic analysis.
- Champion SRE best practices within incident management, focusing on automation, observability, and proactive problem prevention.
- Collaborate with SRE teams to develop and implement incident response playbooks, runbooks, and automation tools to enhance efficiency and effectiveness.
- Contribute to the definition and monitoring of Service Level Objectives (SLOs) and Service Level Indicators (SLIs) related to incident response and system reliability.
- Operational Excellence and Continuous Improvement:
- Analyze incident trends and data to identify systemic issues and areas for improvement in IT systems and processes.
- Develop and implement strategies to reduce Mean Time To Detect (MTTD), Mean Time To Acknowledge (MTTA), and Mean Time To Restore (MTTR).
- Drive initiatives to enhance operational resilience, stability, and performance across critical applications and infrastructure.
- Contribute to the continuous improvement of ETCC incident management policies, procedures, and tools.
- Stakeholder Management and Communication:
- Build strong relationships with key stakeholders across technology and business units.
- Effectively manage expectations and provide transparent communication during periods of service disruption.
- Represent ETCC SRE Operations in various forums and provide updates on incident management performance and initiatives.
- Team Leadership and Development:
- Provide guidance and mentorship to junior incident managers and SRE operations staff.
- Foster a culture of accountability, continuous learning, and incident prevention within the team.
Required Qualifications & Skills:
- Experience:
- Minimum of 10-15 years of experience in IT Operations, Incident Management, or Site Reliability Engineering within a large-scale enterprise environment, preferably in the financial services industry.
- Demonstrated experience in leading major incidents and coordinating cross-functional technical teams under pressure.
- Strong understanding of SRE principles and practices.
- Technical Proficiency:
- Solid understanding of IT infrastructure (servers, storage, networking), cloud platforms (e.g., AWS, Azure, GCP), and enterprise applications.
- Familiarity with incident management tools (e.g., ITSM, Remedy, ServiceNow, PagerDuty) and monitoring tools (e.g., Grafana, Splunk, Dynatrace, Elk).
- Experience with scripting and automation (e.g., Python, Shell) is a plus and added advantage.
- Experience in insurance, banking or regulated financial services environment (mandatory)
- Leadership & Communication:
- Excellent leadership, communication, and interpersonal skills, with the ability to influence and collaborate effectively at all levels.
- Proven ability to remain calm and decisive during critical incidents, with strong problem-solving capabilities.
- Exceptional written and verbal communication skills for technical and non-technical audiences.
- Certifications (Good to Have):
- ITIL V3/V4 Foundation or higher certification.
- Relevant SRE or DevOps certifications.
- Proficiency in using GenAI and Microsoft office
Location:
Hyderabad - DTI Skyview SEZJob:
TechnologySchedule:
RegularEmployee Status:
Full time