Senior Lead Software Engineer - Agentic AI SRE Platform
Are you ready to shape the future of AI-driven reliability engineering? At JPMorganChase, we're building intelligent systems that don't just monitor — they reason, analyze, and act. This is your opportunity to lead at the forefront of agentic AI, designing autonomous systems that transform how we deliver resilient, scalable technology at enterprise scale.
As a Senior Lead Software Engineer at JPMorganChase within the AI/ML Data Platforms organization, you will lead the architecture and development of production-grade agentic AI and multi-agent systems on the Reliability Engineering team. You will design autonomous agents that integrate across telemetry pipelines, application stacks, and cloud infrastructure — driving engineering productivity, system reliability, and operational excellence. This role offers the chance to pioneer how intelligent automation reshapes platform reliability for one of the world's leading financial institutions.
Job responsibilities
- Design and develop enterprise AI systems that enhance engineering productivity, system reliability, and operational excellence
- Architect autonomous agents — not just integrating workflows, but AI systems that reason, analyze, and act across system and application reliability
- Architect and govern agentic system components including orchestration, memory/state management, workflow control, and human-in-the-loop patterns where needed
- Own the agentic architecture non-functional requirements posture end-to-end: observability, security, guardrails, and evaluation
- Provision and manage scalable AI infrastructure, deployment pipelines, and cost optimization
- Execute creative software solutions, design, development, and technical troubleshooting with the ability to think beyond routine or conventional approaches to build solutions or break down technical problems
- Develop secure, high-quality production code, and review and debug code written by others
- Identify opportunities to eliminate or automate remediation of recurring issues to improve overall operational stability of software applications and systems
- Lead evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture
- Contribute to a team culture of diversity, opportunity, inclusion, and respect
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and 5+ years applied experience
- Advanced proficiency in programming with Python
- Proficiency in all aspects of the Software Development Life Cycle and in automation and continuous delivery methods
- Hands-on experience delivering agentic AI to production (not just LLM wrappers or prompt-to-workflow integrations)
- Experience with AI agent frameworks (e.g., LangGraph, CrewAI), runtimes, orchestration, evaluation/benchmarking systems, and model selection policies
- Experience developing production generative AI applications at scale with experimentation, model quality measurement, retrieval-augmented generation, and vector search
- Strong cloud infrastructure experience (e.g., AWS) and infrastructure-as-code (e.g., Terraform)
- Experience building distributed systems or high-scale systems where reliability matters
Preferred qualifications, capabilities, and skills
- Previous experience working within infrastructure, platform engineering, or reliability engineering domains
- Knowledge of site reliability engineering concepts, including Service Level Objectives/Indicators, error budgets, incident management, resilience patterns, and observability
- Previous experience as a machine learning software engineer in a dynamic technology company or startup.
#CTC