Site Reliability Engineer, Data Protection
Our Purpose
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Site Reliability Engineer, Data ProtectionMastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital economy that benefits everyone, everywhere by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential.Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. With connections across more than 210 countries and territories, we are building a sustainable world that unlocks priceless possibilities for all.
Overview
Mastercard’s enterprise storage team is looking for a senior site reliability engineering to help advance our SRE capabilities across our data protection platform with a focus on the Commvault backup suite, CHEF automation, AI implementation, and backup storage technologies from VAST Data, NetApp, Oracle ZFS, and Dell/EMC Data Domain.
• This individual will work closely with technology support teams and application teams to build monitoring and automation solutions to improve the application and infrastructure availability.
• Key questions for viable candidates:
1. Have you resolved a complex application availability issue with monitoring and automation?
2. Are you a self-starter with minimal direction and guidance?
3. Have you ever been part of a team with diverse skills and experience located in different geographical locations/time zones?
Responsibilities
• Provide ongoing support, implementation, and continuous improvement of the Commvault data protection environment to ensure operational stability, performance, and alignment with business requirements.
• Represent enterprise storage team in project meetings and provide advice, status, training, and technical support
• Administer, support, and maintain enterprise Monitoring tools in a multi-tier storage environment
• Build Automation solutions using available tool sets and scripting languages (Ansible, Bitbucket, CHEF, Jenkins)
• Maintain design and support documents for all built solutions and processes
• Troubleshoots networking, Unix/Linux systems, and applications to identify and correct malfunctions and other operational problems utilizing associated Linux and UNIX command line and management tools.
• Work with our customers to understand the monitoring and automation requirements and implement solutions using available tool sets and scripting languages
• As new technologies emerge and impact our environment, learn about these technologies very quickly and resolve any problems involved in integrating new technologies.
• Maintains a broad knowledge of state-of-the-art technology, equipment, and/or systems.
• Thorough, adhering to critical processes even under stress
• Support business disaster recovery procedures for assigned areas of responsibility.
• Accurately document duties and procedures to aid the department in cross-training and absentee coverage
• Work with technical engineering teams to manage and improve processes
All about you:
• Fluent English communication skills, both written and verbal, with the ability to effectively collaborate and influence stakeholders across global teams and executive audiences.
• Advanced expertise with the Commvault Data Protection Suite, including architecture, administration, optimization, and enterprise-scale data protection strategies.
• Deep technical knowledge of Linux/Unix and Windows operating systems, with extensive experience managing and supporting enterprise storage platforms, including Oracle ZFS, Data Domain, VAST Data, and NetApp.
• Experience with enterprise storage infrastructures, including SAN environments, HPE and NetApp storage arrays, and Brocade Fibre Channel networking technologies.
• Advanced experience designing, implementing, and operating enterprise monitoring and observability solutions using platforms such as Grafana, Prometheus, and related technologies.
• Strong understanding of network architecture, protocols, security principles, and infrastructure design, with the ability to troubleshoot complex cross-platform issues.
• Proficiency in automation, scripting, and software development using Python, Bash, and other modern programming languages to improve operational efficiency and platform reliability.
• Proven ability to architect and implement automation solutions that reduce operational toil, improve scalability, and enhance service resiliency.
• Exceptional analytical, troubleshooting, and problem-solving skills, with the ability to quickly diagnose and resolve highly complex technical challenges.
• Strong understanding of Site Reliability Engineering principles, including observability, incident management, root cause analysis, capacity planning, resiliency engineering, disaster recovery, and operational excellence.
• Demonstrated ability to lead multiple strategic initiatives simultaneously while effectively prioritizing competing business and technology demands.
• Excellent communication, presentation, documentation, stakeholder management, and project leadership skills.
• Experience working in large-scale enterprise environments with mission-critical systems requiring high availability, reliability, and performance.
• Self-driven, highly adaptable, and continuously focused on learning emerging technologies and expanding expertise across adjacent technical domains.
• Proven ability to influence technical direction across multiple engineering teams and drive adoption of best practices, standards, and reliability-focused initiatives.
• Strong organizational and time management skills, with the ability to operate effectively in fast-paced environments while maintaining a high level of quality and attention to detail.
• Experience mentoring engineers and providing technical leadership across cross-functional teams in support of enterprise-wide reliability and infrastructure modernization initiatives.
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
Abide by Mastercard’s security policies and practices;
Ensure the confidentiality and integrity of the information being accessed;
Report any suspected information security violation or breach, and
Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.