Staff-Software Engineer

Phoenix, AZFull-timePosted Aug 6, 2026

At American Express, we are transforming how software is built, deployed, and operated at enterprise scale. The Site Reliability Engineering (SRE) organization is at the center of this transformation, building the engineering capabilities that enable secure software delivery, resilient platforms, intelligent operations, and exceptional developer experiences.

As a Staff Engineer, you will be a senior technical leader responsible for advancing the reliability, scalability, and operational excellence of critical technology platforms. You will work across engineering organizations to solve complex production challenges through software engineering, platform innovation, automation, and AI-powered operational capabilities.

This is a highly influential individual contributor role where success is measured by the systems you build, the engineering practices you shape, and the enterprise impact you create.

Role Summary

As a Staff Engineer within the Site Reliability Engineering organization, you will lead the evolution of enterprise reliability engineering by building scalable platform capabilities that improve system resilience, engineering productivity, and operational efficiency.

You will partner closely with Platform Engineering, Infrastructure Engineering, Security Engineering, Application Development, and Enterprise Architecture teams to modernize how software is delivered, observed, secured, and operated across the enterprise.

This role requires deep expertise in distributed systems, cloud infrastructure, software engineering, DevSecOps, observability, automation, and AI-enabled operations. You will drive engineering solutions that reduce operational complexity, improve service reliability, eliminate manual toil, and accelerate the organization's journey toward autonomous operations.

Key Responsibilities

Lead Enterprise Reliability Engineering

Drive the technical direction for reliability engineering across critical enterprise platforms.

You will:

  • Define and evolve enterprise-wide Site Reliability Engineering practices and standards.
  • Champion Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to improve service reliability and engineering accountability.
  • Establish engineering patterns for high availability, scalability, resiliency, and disaster recovery.
  • Lead architecture and design reviews for mission-critical platforms and distributed systems.
  • Partner with engineering teams to improve production readiness, operational maturity, and service resilience.
  • Drive adoption of proactive reliability engineering practices, including capacity planning, performance optimization, failure testing, and resilience validation.
  • Build Software That Improves Operations
  • Apply software engineering to eliminate operational complexity and improve engineering productivity.

  • You will:

  • Design and develop reusable platforms, frameworks, and automation that reduce operational toil.
  • Build engineering capabilities that simplify production operations and enable self-service experiences.
  • Develop reference implementations and engineering libraries adopted across multiple organizations.
  • Improve deployment safety through automated validation, policy enforcement, and release controls.
  • Leverage modern programming languages and cloud-native technologies to solve complex operational problems at scale.

    Modernize Software Delivery & Engineering Platforms

  • Enable secure, reliable, and efficient software delivery across the enterprise.

  • You will:

  • Drive maturity of enterprise CI/CD platforms through standardized pipelines and reusable engineering capabilities.
  • Embed security, reliability, testing, and compliance directly into software delivery workflows.
  • Advance Platform Engineering and Internal Developer Platforms (IDPs) that improve developer experience and engineering velocity.
  • Improve deployment reliability using progressive delivery, automated verification, rollback strategies, and policy-as-code.
  • Partner with application teams to simplify software delivery while improving production stability.

    Advance Intelligent Operations & Autonomous Engineering

  • Lead the evolution of modern operations through AI-driven engineering and intelligent automation.

  • You will:

  • Build next-generation operational capabilities using AI, machine learning, and intelligent automation.
  • Integrate Generative AI and agentic AI into incident management, diagnostics, troubleshooting, and engineering workflows.
  • Design self-healing operational capabilities that automate detection, analysis, and remediation of production issues.
  • Improve operational decision-making through predictive analytics, anomaly detection, and intelligent event correlation.
  • Advance the organization's journey toward autonomous operations while maintaining governance and operational safety.

    Improve Production Security & Operational Resilience

  • Partner with Security Engineering to improve the resilience of production systems through secure engineering practices.

  • Integrate automated vulnerability detection and remediation into software delivery and operational workflows.

  • Build scalable engineering solutions that reduce vulnerability remediation time and improve enterprise security posture.

  • Advance secure-by-default engineering practices across infrastructure and application platforms.

  • Improve software supply chain security through automated controls and continuous compliance.

  • Reduce operational risk through engineering automation and policy-driven governance.

    Drive Enterprise Automation

  • Reduce operational toil through intelligent automation and reusable engineering solutions.

  • Identify repetitive operational activities and replace them with scalable automation.

  • Build reusable automation frameworks supporting infrastructure, application operations, compliance, and engineering workflows.

  • Standardize operational runbooks through workflow orchestration and policy-driven automation.

  • Improve engineering efficiency by enabling self-service operational capabilities.

  • Measure automation effectiveness through reductions in manual effort, operational risk, and incident resolution time.

    Provide Technical Leadership

  • As a Staff Engineer, your influence extends beyond any single team.

  • Lead enterprise-wide engineering initiatives spanning multiple organizations.

  • Influence architectural decisions through technical expertise and engineering excellence.

  • Mentor senior engineers and technical leaders across the organization.

  • Prototype emerging technologies and evaluate their applicability to enterprise engineering challenges.

  • Drive technical alignment through collaboration, technical design reviews, and engineering standards.

  • Foster a culture of continuous improvement, operational excellence, and engineering innovation.

    What Success Looks Like

  • Success in this role will be measured through tangible engineering outcomes, including:

  • Improved service availability, reliability, and platform resilience.

  • Increased adoption of SLO-driven engineering and operational maturity practices.

  • Reduced Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).

  • Increased deployment frequency with improved change success rates.

  • Reduced operational toil through automation and self-service engineering capabilities.

  • Reduced enterprise vulnerability remediation cycle times through engineering automation.

  • Increased adoption of enterprise engineering platforms and reusable automation frameworks.

  • Accelerated adoption of AI-powered operational capabilities and autonomous remediation workflows.

  • Measurable improvements in developer productivity and engineering efficiency.

Minimum Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related technical discipline, or equivalent practical experience.
  • 12+ years of experience in software engineering, Site Reliability Engineering, cloud infrastructure, platform engineering, or distributed systems.
  • Demonstrated experience designing and operating highly available, large-scale distributed systems.
  • Strong software engineering background with proficiency in one or more modern programming languages such as Java, Go, Python, or TypeScript.
  • Deep understanding of cloud-native architectures, Kubernetes, Infrastructure as Code, CI/CD, observability, and engineering automation.
  • Experience applying SRE principles, including SLOs, error budgets, incident management, production readiness, and operational excellence.
  • Proven ability to influence technical direction across multiple engineering organizations.

Preferred Qualifications

Experience with several of the following technologies and practices:

Cloud & Infrastructure

  • AWS
  • Kubernetes/OpenShift
  • Docker
  • Terraform
  • CloudFormation
  • Ansible

Platform Engineering

  • GitHub Actions
  • Jenkins
  • GitLab CI
  • Argo CD

Observability & Reliability

  • OpenTelemetry
  • Prometheus
  • Grafana
  • Splunk
  • Dynatrace
  • Elastic
  • Chaos Engineering
  • Capacity Engineering
  • Performance Engineering

Software Engineering

  • Java
  • Go
  • Python
  • Event-driven architectures
  • Kafka
  • Distributed systems design

AI & Automation

  • Generative AI
  • Agentic AI
  • AIOps
  • Workflow orchestration
  • Intelligent automation platforms
  • AI-assisted engineering tools

Leadership Expectations

Successful Staff Engineers at American Express:

  • Solve complex engineering problems through software, automation, and scalable platform solutions.
  • Balance innovation with operational excellence, resilience, and security.
  • Influence engineering direction through technical expertise rather than organizational authority.
  • Remain hands-on by designing systems, reviewing critical implementations, and developing prototypes.
  • Build consensus across engineering organizations through collaboration and technical credibility.
  • Mentor engineers and cultivate a culture of learning, ownership, and engineering excellence.
  • Continuously evaluate emerging technologies that improve reliability, developer productivity, and operational efficiency.

Why Join Us?

As a Staff Engineer in the Site Reliability Engineering organization, you will help shape the future of engineering at American Express. Your work will influence how thousands of engineers build, deploy, secure, observe, and operate software across one of the world's largest financial technology platforms.

You will have the opportunity to build foundational engineering capabilities, modernize enterprise software delivery, advance AI-powered operations, and drive the next generation of autonomous engineering—all while improving the reliability and resilience of services that millions of customers depend on every day.

Employment eligibility to work with American Express in the United States is required as the company will not pursue visa sponsorship for these positions.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free