Lead Software Engineer - Java , Kubernetics, Infrastructure
We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.
As a Lead Software Engineer at JPMorganChase, within Delivery Platforms you will be a technical leader and hands-on builder of the firm's composable test environment platform — enabling developers across a 10,000-engineer organization to spin up isolated, production-representative virtual environments per pull request, run integration tests independently, and merge without contention on shared environments. You will operate across hybrid infrastructure (on-premises Kubernetes and AWS), solve complex problems in traffic routing, dependency graph resolution, and stateful service provisioning, and deliver a platform that thousands of engineers rely on daily.
Job Responsibilities
- Designs and implement core subsystems of the composable environment platform — including the provisioning control plane, lifecycle management engine, resource scheduling, and automatic teardown with time-to-live enforcement
- Builds and own the traffic routing and isolation layer that directs requests to the correct service versions within virtual environment sessions using header-based routing, service mesh integration, and service discovery — ensuring zero cross-session contamination.
- Implements dependency graph resolution logic that determines which services are instantiated live, reused from baseline, or virtualized/stubbed for a given test scenario — enabling composable environments at 1,000+ service scale
- Develops and maintain the self-service interfaces (CLI and API) that enable application teams to declare environment needs, provision in minutes, execute tests, and tear down — with clear contracts, versioning, and backward compatibility
- Builds automated provisioning for stateful test infrastructure across heterogeneous data stores (Oracle, Kafka, Cassandra, MongoDB, CockroachDB, DynamoDB, PostgreSQL), including snapshot/restore, schema versioning, and synthetic data seeding
- Implements CI/CD pipeline integrations — environments created on PR open, tests executed, results reported, environments destroyed on merge/close — across Jenkins, Spinnaker, GitHub Actions, and Argo CD
- Builds observability, operational dashboards, and alerting for the platform itself — tracking time-to-environment, provisioning success rates, resource utilization, cost attribution, and isolation correctness
- Ensures the platform operates reliably across hybrid infrastructure (on-premises Kubernetes and EKS) with consistent provisioning semantics, cross-network connectivity, and acceptable spin-up latency regardless of hosting tier. Mentors and provide technical guidance to staff engineers on the team, conduct design reviews, and establish coding standards and testing practices for platform code
- Architects and governs agentic AI-enabled engineering workflows (using enterprise-authorized tools within the work environment) to improve delivery speed, code quality, and operational outcomes at scale (e.g., AI-driven PR review assistance, test generation/maintenance, release readiness checks, incident triage and root-cause acceleration), while defining guardrails for validation, security, resiliency, and reuse across teams
- Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
Required Qualifications, Capabilities, and Skills
- 12+ years of professional software engineering experience, with significant hands-on experience building internal platforms, infrastructure automation, or developer tooling consumed by large engineering organizations
- Strong hands-on proficiency in Go, Rust, or Java for building production-grade platform services — including API design, concurrency, fault tolerance, and operational reliability at scale
- Deep experience with Kubernetes (on-premises and EKS) — including namespace management, custom controllers/operators, resource scheduling, multi-tenancy patterns, and networking (CNI, ingress, service mesh)
- Hands-on experience with traffic routing and isolation patterns — header propagation, service mesh traffic control (Istio, Linkerd, or equivalent), service discovery, and request-scoped environment resolution
- Strong infrastructure-as-code proficiency (Terraform, Ansible) with disciplined change control, state management, and repeatable provisioning across hybrid cloud and on-premises environments
- Proven experience provisioning and managing stateful services in automated fashion — relational databases (Oracle, PostgreSQL), document stores (MongoDB, CockroachDB), message brokers (Kafka), and cloud-native services (DynamoDB)
- Experience building CI/CD pipeline extensions and integrations — embedding custom stages, webhooks, and lifecycle hooks into Jenkins, Spinnaker, GitHub Actions, or Argo CD
- Strong systems thinking — designing for composability, multi-tenancy, horizontal scalability, graceful degradation, and operational resilience in distributed systems
- Ability to mentor engineers, lead design reviews, and drive technical quality through code review, testing standards, and documentation practices
- Demonstrated experience designing and leading adoption of agentic AI-enabled development practices (using enterprise-authorized tools within the work environment) across teams, including setting standards for human-in-the-loop validation, auditability/traceability of changes, and secure handling of sensitive data
- Strong understanding of responsible AI use and control expectations in engineering workflows, including security/resiliency implications, data sensitivity, and risk-based governance; ability to influence senior technical leaders on safe scaling patterns and reuse
Preferred Qualifications, Capabilities, and Skills
- Experience building environment-as-a-service platforms, composable/virtual environment systems, or namespace-per-PR isolation patterns at scale
- Experience implementing test data strategies at scale — masked production snapshots, on-demand restore, synthetic data generation, and repeatable dataset management across diverse data technologies
- Experience with FinOps practices — cost attribution, chargeback modeling, idle resource detection, and automated reclamation policies
- Experience applying AI/ML techniques to infrastructure problems — predictive scaling, intelligent composition, or automated dependency resolution