Site Reliability Engineer II
Play a key role in ensuring system reliability at one of the world’s most iconic and largest financial institutions.
As a Site Reliability Engineer II at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will use technology to solve business problems and leverage software engineering best practices as we strive towards excellence. This role often works independently to execute small to medium projects, but you’ll also have the opportunity to collaborate with cross functional teams to continually improve your level of knowledge about JPMorgan Chase’s business and relevant technologies.
- Executes small to medium projects independently with initial direction and graduates to designing and delivering projects independently
- Leverages technology to solve business problems by writing high quality, maintainable, and robust code following best practices in software engineering
- Uses enterprise-authorized AI capabilities within the work environment to speed up incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Participates in triaging, examining, diagnosing, and resolving incidents and works with others to solve problems at their root
- Recognizes toil within the role and proactively works towards eliminating it through systems engineering or updating application code
- Understands observability patterns and strives to implement and improve service level indicators, objectives monitoring, and alerting solutions for optimal transparency and analysis
- Applies enterprise-authorized AI capabilities within the work environment to identify recurring toil and reliability risks from operational signals, prioritizing reuse-first improvements and measurable SLO outcomes.
- A. Formal training or certification on software engineering concepts and 2+ years applied experience ( NAMR/APAC – India/ LATAM/ Hong Kong)
B. Formal training or certification on software engineering concepts and advanced applied experience (EMEA/LATAM-Brazil)
C. Singapore follow local country guidance - Ability to code in at least one programming language
- Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows (e.g., troubleshooting support and runbook drafting) with strong validation habits and awareness of data sensitivity.
- Ability to assess AI-assisted operational recommendations for correctness and risk, and apply appropriate controls to maintain resiliency, security, and auditability.
- Experience maintaining a cloud-based infrastructure
- Familiar with site reliability concepts, principles, and practices
- Familiar with observability such as white and black box monitoring, service level objective alerting, and telemetry collection
- Familiarity with containers or a common server OS such as Linux and Windows
Emerging knowledge of continuous integration and continuous delivery practices and related tooling