Technology Support Lead
Join our dynamic team to innovate and refine technology operations, impacting the core of our business services.
As a Technology Support Lead DevOps Engineer at JPMorganChase within Global Banking Technology (Sales & Marketing domain applications and Experience Layer) team, you will lead operational excellence for critical client-facing and internal platforms. You will partner with product, engineering, and platform teams to improve availability, performance, resiliency, and delivery outcomes through automation, enhanced observability, disciplined incident/problem management, and strong release execution. Your role will require to support AWS-based services and aligned platforms (e.g., Gaia Kubernetes Platform (GKP), GAP) and will require flexibility to support production releases, platform upgrades, and incident response windows.
Job Responsibilities
- Own and drive a DevOps strategy for Sales & Marketing domain applications and Experience Layer services, delivering measurable improvements in availability, latency, deployment success rates, and operational health.
- Define and operationalize SLIs/SLOs and error budgets; implement reliability scorecards and service health reporting for senior stakeholders, and establish production standards for observability (metrics, logs, traces), alerting, dashboards, and on-call readiness; reduce alert noise and improve actionable detection.
- Identify, implement, and continuously tune alerts (symptom-based and SLO-aligned) to improve early detection, reduce false positives, and accelerate triage.
- Lead incident management for production events: command/coordination, stakeholder communications, post-incident reviews, and corrective action tracking to closure and drive toil reduction through automation (Python) across operational workflows, environment checks, runbook automation, and self-healing patterns.
- Own release execution and release risk management for supported services, including release readiness validation, run-of-show/cutover coordination, go/no-go facilitation, post-release verification, and rollback/backout planning.
- Partner with application and platform teams to standardize release processes (checklists, approvals, validations, release communications, and evidence capture) to improve change success rates and reduce production incidents, mentor and coach DevOps/production engineers; contribute to hiring, onboarding, and setting team standards and best practices.
- Partner with engineering teams to embed resiliency-by-design practices: capacity planning, performance testing, dependency mapping, failover validation, and controlled experiments (where appropriate), and collaborate with cloud/platform teams to ensure robust operations across AWS and aligned enterprise platforms (including GKP and GAP as applicable).
- Drive periodic platform and cluster upgrades (e.g., Kubernetes lifecycle upgrades on GKP and AWS EKS upgrades), including upgrade planning, risk assessment, testing/validation, cutover execution, and post-upgrade stabilization. Lead and coordinate data platform upgrades and maintenance activities, including GOS upgrades and Cassandra upgrades (version lifecycle management, compatibility checks, performance/availability validation, rollback planning).
- Enhance service observability by improving end-to-end telemetry coverage (Dynatrace/Splunk/Grafana), strengthening dashboards for critical user journeys, and ensuring logs/metrics/traces support fast root cause analysis.
- Leads team adoption of enterprise-authorized AI capabilities within the work environment to improve incident triage speed and consistency (e.g., synthesizing operational signals into prioritized actions), with human-in-the-loop validation and appropriate handling of sensitive data.
- Applies reuse-first, AI-assisted practices across incident/problem/change routines to identify recurring interruption patterns and validate remediation actions aligned to resiliency and security expectations.
Required Qualifications, Capabilities, and Skills
- 5+ years of experience or equivalent expertise troubleshooting, resolving, and maintaining information technology services
- Strong experience in DevOps / Production Engineering / Platform Operations with experience in supporting large-scale, distributed, high-availability systems.
- Demonstrated leadership in driving outcomes across multiple teams (engineering, product, operations) with clear metrics and accountability.
- AWS operations experience, including reliability patterns for cloud-native services (networking, scaling, resiliency, monitoring) with strong proficiency in Python for automation and operational tooling (operational utilities, APIs, runbook automation, data collection).
- Hands-on experience driving release execution in production environments, including dependency coordination, release readiness, cutover planning, and post-release validation.
- Experience planning and executing platform/cluster upgrades with minimal disruption, including stakeholder coordination and validation practices with working knowledge of operating and upgrading distributed data systems, including GOS and Cassandra (and associated operational risk considerations).
- Deep knowledge of observability practices: telemetry design, alert strategy/tuning, incident detection, and MTTR reduction with proven incident leadership experience (severity management, communications, postmortems) and ability to drive remediation and long-term fixes with strong fundamentals in Linux and networking (HTTP/TLS, DNS, load balancing, latency troubleshooting).
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support production operations workflows with strong validation habits and awareness of data sensitivity.
- Ability to review and validate AI-assisted incident recommendations before action, escalating when uncertain and ensuring outcomes align to operational, security, and auditability expectations.
Preferred Qualifications, Capabilities, and Skills
- Experience supporting client-facing experience layers (web/mobile/API gateways, edge/perimeter patterns, high-traffic front-door services).
- Familiarity with internal platforms/accelerators such as GKP and GAP (or equivalent enterprise platform services), and integrating DevOps standards across them.
- Experience with containers/orchestration (e.g., Kubernetes) and service-to-service reliability patterns.
- Infrastructure-as-Code and configuration automation experience.
Track record of eliminating manual release/ops steps through automation and improving developer/support experience via standardized tooling and guardrails.