Software Engineer III - DevOps/ML Ops AWS Streaming
You’re ready to gain the skills and experience needed to grow within your role and advance your career — and we have the perfect software engineering opportunity for you.
As a Software Engineer III at JPMorganChase within the authentication, identity verification, and fraud detection engineering team (DevOps and Platform Engineer), you will build, operate, and scale production platforms that enable secure, real-time machine learning services. You will partner closely across engineering, data science, and security to deliver reliable, observable, and resilient services that support critical customer and firm outcomes.
Job Responsibilities
- Provision and manage Amazon Web Services Kubernetes and container environments, including networking, access controls, cluster configuration, and autoscaling, using Terraform to support consistent, repeatable delivery.
- Build reusable infrastructure-as-code modules, define standards, and reduce configuration drift through strong environment hygiene and remote state management practices.
- Enable end-to-end machine learning operations workflows across build, validation, packaging, deployment, monitoring, and retraining to support production model lifecycle needs.
- Implement robust deployment and release patterns for machine learning-enabled services, including versioning, progressive delivery, and rollback strategies aligned to reliability goals.
- Operate and troubleshoot Apache Kafka streaming integrations, addressing throughput, latency, resiliency, and consumer lag to sustain real-time decisioning workloads.
- Deploy, operate, and scale Apache Flink on Kubernetes, managing job lifecycle, state, checkpoints, upgrades, recovery, and performance tuning.
- Deploy, maintain, and scale Ray Serve on Kubernetes to provide low-latency inference and service orchestration, improving resource efficiency and runtime stability.
- Implement and continuously improve observability (logs, metrics, tracing), dashboards, and alerting, and lead incident response and corrective actions to reduce mean time to recovery.
- Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity across complex deliverables (e.g., code generation/refactoring, unit test creation, documentation), while validating outputs through peer review, automated testing, and secure coding standards; contributes learnings and reusable patterns to improve broader team effectiveness.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
Required Qualifications, Capabilities, and Skills
- Formal training or certification on software engineering concepts and 3+ years applied experience
- Proven experience in DevOps, platform engineering, site reliability engineering, or machine learning operations supporting production services.
- Hands-on experience with Amazon Web Services and Kubernetes (Amazon Elastic Kubernetes Service), with ability to provision, run, and troubleshoot production clusters.
- Strong experience with Terraform, including modules, multi-environment patterns, remote state, and continuous integration/continuous delivery integration to reduce drift.
- Experience operating Kafka-based streaming systems and resolving production pipeline issues related to performance, reliability, and resiliency.
- Experience deploying and operating distributed compute platforms on Kubernetes, including Apache Flink.
- Strong Linux and networking fundamentals with scripting and automation skills (Python and/or Bash) and a clear operational ownership mindset.
- Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, test creation, troubleshooting, or documentation) with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security.
- Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices.
Preferred Qualifications, Capabilities, and Skills
- Experience with GitOps and Kubernetes packaging tooling (for example, Argo CD, Flux, Helm, or Kustomize).
- Experience with observability stacks such as Prometheus and Grafana, OpenTelemetry, Elasticsearch or OpenSearch, Splunk, or Datadog.
- Experience with Kubernetes policy and security controls (for example, Open Policy Agent, Gatekeeper, or Kyverno) and secrets tooling (for example, Vault or external secrets controllers).
- Experience with high-sensitivity domains such as authentication, identity verification, fraud, risk, or other regulated financial workloads.
- Performance tuning experience for low-latency inference services and real-time streaming workloads.