Sr. Lead Software Engineer - AIML Platforms
Be an integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch technology products.
As a Sr. Lead Software Engineer at JPMorgan Chase within the AI/ML Platforms team in Corporate Sector, you will design, build, and operate the foundational cloud infrastructure that enables data scientists and machine learning engineers to develop, train, and deploy intelligent solutions across the firm. You will serve as a technical leader, driving platform reliability, scalability, and automation while collaborating with cross-functional teams to solve complex infrastructure challenges. Your work will directly accelerate the firm's AI/ML capabilities, enabling faster experimentation and production-grade deployments that create measurable business impact.
Job responsibilities
- Architect and implement scalable, production-grade infrastructure on AWS using Kubernetes (EKS) to support AI/ML workloads across the enterprise
- Lead the design and development of infrastructure-as-code solutions using Terraform, ensuring repeatable, auditable, and automated environment provisioning
- Drive the installation, configuration, and lifecycle management of AI/ML platform components, ensuring high availability and operational resilience
- Partner with data science, machine learning, and product teams to define platform requirements and translate them into robust engineering solutions
- Establish and enforce engineering best practices, coding standards, and security controls across the platform infrastructure
- Identify and resolve performance bottlenecks, reliability gaps, and scalability constraints within the AI/ML platform ecosystem
- Contribute to the continuous improvement of CI/CD pipelines, enabling faster and safer delivery of platform updates and model deployments
- Mentor and guide junior engineers, fostering a culture of technical excellence, knowledge sharing, and collaborative problem-solving
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and 5+ years applied experience
- Demonstrated expertise with infrastructure-as-code tooling, specifically Terraform, in large-scale cloud environments
- Hands-on experience designing and managing Kubernetes clusters, with specific depth in AWS Elastic Kubernetes Service (EKS)
- Proven ability to architect and operate cloud-native infrastructure on AWS, including compute, networking, storage, and security services
- Experience supporting or building platforms for AI/ML workloads, including model training, serving, or experimentation environments
- Strong understanding of DevOps and platform engineering principles, including CI/CD, observability, and automated testing
- Ability to lead technical initiatives independently, communicate complex concepts clearly to both technical and non-technical stakeholders, and drive solutions from design through production
Preferred qualifications, capabilities, and skills
- Proficiency in Go or Python for automation, tooling development, or platform service implementation
- Experience with MLOps frameworks and tools such as Kubeflow, MLflow, or similar AI/ML lifecycle management platforms
- Familiarity with service mesh technologies, cloud networking patterns, and container security best practices
- Exposure to multi-cloud or hybrid cloud architectures and platform portability strategies