Site Reliability Engineer II
Join a dynamic team at the forefront of technology, where your skills drive innovation and modernize mission-critical systems. Experience career growth and make a meaningful impact.
As a Site Reliability Engineer II at JPMorgan Chase within the Enterprise Technology, Engineering Services and Platform team, you will solve complex business problems with straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications and their associated infrastructure to iteratively improve existing solutions. You are a significant contributor by sharing your knowledge of end-to-end operations, availability, reliability, and scalability of your application or platform. We foster a collaborative environment where your expertise shapes the future of technology.
Job responsibilities
- Guide and assist others in building appropriate level designs and gaining consensus from peers
- Collaborate with software engineers and teams to design and implement deployment approaches using automated CI/CD pipelines
- Design, develop, test, and implement availability, reliability, and scalability solutions in applications
- Implement infrastructure, configuration, and network as code for applications and platforms
- Collaborate with technical experts, stakeholders, and team members to resolve complex problems
- Understand service level indicators and utilize service level objectives to proactively resolve issues
- Support the adoption of site reliability engineering best practices within your team
- Accelerate delivery and improve operational rigor through AI-assisted engineering adoption
- Provide 24/7 production support for business-critical applications
- Uses enterprise-authorized AI capabilities within the work environment to speed up incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Applies enterprise-authorized AI capabilities within the work environment to identify recurring toil and reliability risks from operational signals, prioritizing reuse-first improvements and measurable SLO outcomes.
Required qualifications, capabilities and skills
- Formal training or certification on software engineering concepts and 2+ years applied experience
- Proficient in site reliability culture and principles, and familiarity with implementation within applications or platforms
- Proficient in at least one programming language such as Python, Java/Spring Boot, or .Net
- Experience in observability including monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Elastic, Splunk
- Experience with CI/CD tools like Jenkins, GitLab, or Terraform
- Experience with event streaming platforms like Kafka
- Experience with AI-assisted tools like GitHub Copilot, Claude
- Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows (e.g., troubleshooting support and runbook drafting) with strong validation habits and awareness of data sensitivity.
- Ability to assess AI-assisted operational recommendations for correctness and risk, and apply appropriate controls to maintain resiliency, security, and auditability.
Preferred qualifications, capabilities and skills
- Deep understanding of TCP/IP, DNS, load balancing, firewalls, and VPN technologies
- Experience tuning Linux performance and troubleshooting system-level issues
- Certifications: AWS Certified SysOps Administrator or Professional, Certified Kubernetes Administrator (CKA), or equivalent
- Familiarity with container and orchestration technologies such as ECS, Kubernetes, and Docker