Sr. Staff Cloud Platform Engineer - Security
Atlanta, GA · Sunrise, FL · San Francisco, CA · Seattle, WA · Lowell, MAPosted Jun 29, 2026
Architectural Leadership & Consulting: Act as the primary resilience advisor to multiple distributed product and enterprise teams, guiding them on best practices for building high availability (HA) and redundancy into their SaaS applications. Resilient Cloud Design: Design and recommend fast-failover solutions and highly available infrastructure primarily on Google Cloud Platform (GCP), while also providing oversight for workloads in Azure and AWS. Infrastructure Validation: Leverage your strong background in Infrastructure as Code (IaC) to review, validate, and guide the implementation efforts of engineering teams. Container & Compute Resilience: Design redundancy strategies for workloads running on Google Kubernetes Engine (GKE) and virtual machines, ensuring self-healing deployments. Cross-Functional Collaboration: Partner closely with DevOps, SRE, and Product Engineering teams to champion resilience engineering principles, chaos testing, and failover validations across tier-0 mission-critical systems. Cloud Platform Expertise: Deep, practical technical knowledge of Google Cloud Platform (GCP) core services, specifically GKE, Compute Engine, and CloudSQL. Familiarity with AWS and Azure is highly desirable. Technical Practitioner Background: Proven past experience as a hands-on engineer who has deployed complex infrastructure. You should understand the implementation details well enough to effectively guide engineering teams. High Availability Architecture: Demonstrated success in architecting active-active or active-passive fast failover mechanisms for high-volume, data-intensive SaaS applications. Database Resilience: Strong understanding of database clustering, replication, and migration strategies (especially migrating legacy RDBMS like MS SQL Server to cloud-native solutions like CloudSQL). Advisory Skills: Excellent communication and consulting skills, with the ability to influence technical teams, explain complex architectural concepts, and foster a culture of resilience without having direct reporting authority over the engineering teams. Resilient Network Services: Practical design experience managing high-availability network topologies, including load balancing, DNS & name resolution, firewalls/gateways, identity/authentication systems, and centralized logging/SIEM. User Session Management: Deep understanding of user session replication, session state persistence, and failover routing strategies in high-traffic, multi-region application architectures. Broad HA Domain Exposure: Familiarity assessing or designing resilience across a comprehensive range of critical SaaS failure domains, such as API gateways, caching layers, messaging/queuing systems, and CI/CD pipelines.