Senior Lead Infrastructure Engineer - Enterprise Technology Data Protection & Recovery

United StatesFull-timePosted Jul 31, 2026

Push the limits of what’s possible with us as an experienced senior member of our product team. 

As a Senior Lead Infrastructure Engineer at JPMorgan Chase within the Enterprise Technology Data Protection & Recovery product line, you will operate at the intersection of infrastructure engineering, product ownership, and risk. You will own the product vision, roadmap, and backlog for the Resiliency Evidence Service (RES), the platform of record for capturing, validating, and reporting resiliency evidence across the firm, while providing hands-on technical leadership to the platform, application, and infrastructure teams that produce and consume that evidence. You will translate control objectives, regulatory expectations, and recovery targets (RTO/RPO) into concrete engineering guidance, evidence contracts, and delivery plans, then partner with global development and infrastructure teams to make those outcomes real.

Required qualifications, capabilities, and skills

  • Own the product vision, roadmap, and quarterly plan for the Resiliency Evidence Service; translate firmwide resiliency, risk, and audit objectives into a prioritized backlog with clear outcomes, success metrics, and release milestones.
  • Act as the single-threaded product owner: groom and prioritize the backlog, run sprint planning and reviews, manage cross-team dependencies, and communicate status, risks, and trade-offs to engineering, product, risk, and executive stakeholders.
  • Provide hands-on infrastructure engineering leadership: define reference architectures, evidence contracts (schemas, APIs, event formats), and integration patterns that partner teams use to emit resiliency evidence into RES.
  • Guide development and infrastructure teams on controls and control objectives from a risk perspective; interpret firmwide standards, recovery objectives, and audit findings into concrete engineering requirements, acceptance criteria, and evidence artifacts.
  • Partner with platform, application, and infrastructure teams to help them design, instrument, and produce the telemetry, attestations, restore proofs, and control evidence consumed by RES; run working sessions, review designs, and unblock adoption.
  •  Steward the end-to-end architecture of RES ingestion, storage, validation, and reporting surfaces; define domain boundaries, service contracts, and cross-service standards, and maintain architectural decision records.
  • Drive resilient service design for RES itself: manage availability SLOs, latency budgets, and error budgets; lead DR tests, restore validation exercises, and chaos/resilience drills for the platform and its dependencies.
  • Champion the secure SDLC across the ecosystem: threat modeling, SAST/DAST integration, dependency and SBOM management, secrets hygiene, encryption in transit and at rest, and robust authentication/authorization patterns for evidence pipelines.
  • Advance observability and SRE practices: instrument metrics, logs, and traces across ingestion and reporting paths; design actionable alerts; author runbooks; lead incident response and blameless postmortems for both RES and the evidence supply chain.
  • Support CI/CD and test quality across the product line: maintain an end-to-end understanding of how RES services, evidence pipelines, and partner integrations fit together and are tested (unit, integration, contract, performance); set expectations for coverage, quality gates, and progressive delivery (canary/blue-green) with rollback plans; review results, triage failures, and guide engineers to the likely root cause when tests or pipelines break.
  • Represent RES with Risk, Controls, Audit, and Compliance partners; produce data-backed status updates, escalation requests, and evidence packages that demonstrate control effectiveness and recovery posture.
  • Support critical environments (Development, QA, Simulation, Production) and lead management of on-call obligations for owned services, ensuring operational stability and timely resolution of incidents.
  • Contribute to the engineering community as an advocate of firmwide frameworks, tools, and best practices; add to team culture by fostering diversity, inclusion, and respect.

 

Required qualifications, capabilities, and skills

  • 5+ years’ experience in infrastructure, platform, or software engineering, with a proven track record leading cross-team delivery of production platforms in a large, regulated environment.
  • Demonstrated product ownership experience: authoring and maintaining a roadmap, running a prioritized backlog, leading sprint ceremonies, and managing stakeholder expectations across engineering, product, and risk audiences.
  • Deep knowledge of resiliency engineering and disaster recovery objectives with hands-on practice running DR tests, restore validation, and chaos/resilience exercises.
  • Strong understanding of technology risk and control frameworks: ability to interpret control objectives, regulatory expectations, and audit findings and translate them into concrete engineering requirements, evidence artifacts, and acceptance criteria.
  • Hands-on infrastructure engineering experience: Linux, containers and orchestration (Kubernetes), infrastructure as code (Terraform or equivalent), multi-AZ/region architectures, and cost-efficient scalability; hands-on experience with AWS and awareness of other cloud providers.
  • Working proficiency in at least one general-purpose programming language (e.g., Python, Java) sufficient to understand and design APIs, define integration tooling, prototype evidence pipelines, and review partner-team code.
  • Data and integration fluency: API design (REST/gRPC), event-driven patterns, schema evolution and compatibility, and practical experience with streaming or messaging platforms (Kafka or equivalent).
  • Observability discipline: instrumentation of metrics/logs/traces, alerting design, runbooks, incident management, and blameless postmortems.
  • Secure SDLC expertise: threat modeling, SAST/DAST integration, dependency risk/SBOM management, secrets handling, encryption standards, and robust authentication/authorization.
  • DevOps/SDLC: familiarity with build tools, CI/CD systems, and version control (Git) with automated pipeline governance.
  • Excellent written and verbal communication; able to move fluidly between deep technical review with engineers and outcome-oriented conversations with senior stakeholders.
  • Ability to independently tackle complex design, delivery, and prioritization problems with minimal oversight.

 

Preferred qualifications, capabilities, and skills

  • Prior ownership of a firmwide platform, control tool, or shared service consumed by many partner teams.
  • Experience defining evidence, control, or telemetry contracts (schemas, attestations, control mappings) and driving adoption across independent engineering teams.
  • Familiarity with data protection and recovery domains: backup and restore tooling, immutable storage, chain-of-custody controls, and auditability.
  • Experience with service mesh, API gateways, and contract governance at scale.
  • Knowledge of Windows/UNIX/Linux and shell scripting for operational tooling and automation.
  • Familiarity with modern frontend technologies for internal tooling or dashboards.
  • Experience establishing engineering standards across multiple teams and leading technical working groups.
  • Exposure to regulatory and audit regimes applicable to financial services resiliency (e.g., operational resilience, recovery and resolution planning).

 

Leadership and collaboration expectations

  • Act as a trusted technical and product leader across regions, influencing peers and decisionmakers to adopt leading-edge practices in resilience, risk, and platform engineering.
  • Set direction for partner teams by publishing clear evidence contracts, adoption guides, and roadmaps; then partner shoulder-to-shoulder with those teams to help them deliver.
  • Represent the product to senior stakeholders in Risk, Controls, Audit, and executive forums with data-backed narratives on progress, coverage, and residual risk.
  • Mentor engineers at multiple levels, elevate code quality and design rigor, and contribute to internal forums and tech talks to disseminate best practices.
  • Build strong working relationships with platform partners across Product Development, Infrastructure, Operations, Risk/Controls, and other firmwide functions to deliver competitive, scalable, and defensible solutions.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free