Lead Infrastructure Engineer - Linux / Windows
Join Corporate Technology at JPMorganChase, where infrastructure engineers are empowered to build resilient, scalable platforms that power critical business intelligence capabilities across the firm. Here, your work directly enables data-driven decisions at enterprise scale—and you'll have the tools, autonomy, and collaborative environment to make a lasting impact.
As a Lead Infrastructure Engineer at JPMorganChase within Corporate Technology, you will design, build, and operate the hybrid infrastructure supporting enterprise business intelligence platforms across on-premises and cloud environments. You will drive platform reliability, automation, and AI-enabled operational efficiencies that reduce toil, accelerate incident response, and elevate the quality of operational knowledge management. Your contributions will directly influence platform resilience, engineering standards, and the adoption of intelligent automation across the team.
Job responsibilities
- Engineer and sustain hybrid infrastructure environments supporting enterprise analytics platforms across development, test, and production tiers, including compute, storage, networking, load balancing, DNS, certificates, and OS configurations
- Provision and manage infrastructure using Terraform, building reusable modules and standardized patterns that enable consistent, repeatable deployments across environments
- Automate recurring operational tasks—including environment builds, validation checks, certificate rotations, and health checks—to reduce manual toil and improve platform reliability
- Drive platform resiliency through high availability design, backup and restore procedures, disaster recovery planning and testing, and environment standardization
- Implement and maintain monitoring and alerting for infrastructure and platform dependencies, including system metrics, service health, log signals, and network performance
- Lead incident response, root-cause analysis, and corrective action tracking, producing structured documentation that reduces repeat incidents and improves mean time to resolution
- Identify and implement AI and large language model-assisted efficiencies across operational workflows, including incident summarization, runbook generation, alert correlation, and change-plan validation
- Partner with platform administrators and application and data teams to align infrastructure capacity, job scheduling, concurrency, and performance baselines with platform requirements
- Produce and maintain operational documentation, including architecture diagrams, standard operating procedures, troubleshooting runbooks, and support checklists
- Uses enterprise-authorized AI capabilities within the work environment to accelerate infrastructure analysis and design documentation, validating outputs and handling operational data according to sensitivity and security requirements.
- Applies reuse-first, AI-assisted practices within delivery and automation routines to identify recurring issues and validate remediation options, ensuring changes are traceable/auditable and aligned to resiliency and security expectations.
Required qualifications, capabilities, and skills
- Formal training or certification on infrastructure engineering concepts and 5+ years applied experience
- Demonstrated experience engineering and supporting production infrastructure in hybrid on-premises and cloud environments
- Strong Linux administration and troubleshooting skills, including service health, system logs, resource contention, process management, and connectivity diagnostics
- Hands-on experience with Terraform, including module development, state management, environment separation, and repeatable infrastructure provisioning
- Experience with IT service management processes and tooling, including incident, change, and problem management workflows
- Proven ability to drive automation and reliability improvements with measurable operational outcomes
- Strong production ownership mindset with structured root-cause analysis skills and calm, effective execution under pressure
- Clear written and verbal communication skills, with the ability to engage both technical teams and business stakeholders
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support infrastructure engineering workflows with strong validation habits and awareness of data sensitivity.
- Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations.
Preferred qualifications, capabilities, and skills
- Experience supporting infrastructure for enterprise analytics or business intelligence platforms, including scaling concepts, scheduling and extract window management, and dependency troubleshooting
- Demonstrated experience applying AI or large language model tooling to operational workflows—such as triage, documentation generation, or automation—with an evaluation and continuous improvement mindset
- Scripting and automation experience using Python or Bash for operational tooling and workflow enhancement
- Familiarity with observability tooling, including log aggregation, metrics platforms, and alert tuning
- Knowledge of security fundamentals relevant to infrastructure operations, including certificate lifecycle management, secrets management, least privilege principles, and vulnerability remediation
- Understanding of network fundamentals applicable to enterprise platforms, including load balancers, reverse proxies, firewall rules, DNS, and TLS