Lead Site Reliability Engineer
New York, NY$179k–$226kPosted Apr 7, 2026
This website utilizes technologies such as cookies to enable essential site functionality, as well as for analytics, personalization, and targeted advertising. To learn more, view the following link: Privacy Policy Manage Preferences
See all jobs
Engineering
Lead Site Reliability Engineer
New York City
Alloy is where you belong!
Alloy helps solve the identity risk problem for companies that offer financial products by enabling them to outpace fraud and confidently serve more people around the world. Over 800 of the world’s largest financial institutions and fintechs turn to Alloy to take control of fraud, credit, and compliance risk, and grow with the clearest picture of their customers.
Through our values: Be Bold, Get Scrappy, Collaborate, and Celebrate Our Differences, we are creating a workplace where you can grow, thrive, and belong. See how we’ve been continuously recognized and named one of Inc. Magazine’s Best Workplaces, Forbes America’s Best Startup Employers, Best Fintech to Work for by American Banker, year after year.
Check out our investors and read more about us here.
About the team
Alloy’s Infrastructure Team is a small team (6 engineers) responsible for a large and growing infrastructure footprint: 15+ Kubernetes clusters, 100+ databases, dozens of services, and complex data organization.
Our challenge isn’t just scale—it’s making that scale reliable, secure, and operable with less manual work.
We’re looking for engineers who enjoy turning complex, fragile systems into automated, self-service platforms with strong safety guarantees.
What you'll be doing
Reporting to the Engineering Manager of Infrastructure, you'll:
Design and build systems to automate infrastructure management at scale (provisioning, upgrades, migrations)
Reduce operational toil by turning manual processes into reliable, repeatable workflows
Build internal tooling and platforms that enable safe self-service changes for other engineers
Improve the reliability and resilience of our infrastructure (Kubernetes, databases, services)
Implement and evolve systems for deploying and running applications in Kubernetes
Contribute to architecture decisions across infrastructure, reliability, and security
Write and review production-quality code
Participate in on-call rotations—but focus on building systems that prevent incidents, not just respond to them
Who we’re looking for
10+ years of experience in infrastructure, SRE, or software engineering roles
Strong software engineering skills—you build systems, not just scripts
Experience managing production infrastructure at scale (cloud + containerized systems)
Experience with Infrastructure as Code (e.g., Terraform)
Experience running and troubleshooting distributed systems (Docker/Kubernetes)
Experience with observability and debugging tools (Datadog, CloudWatch, ELK/EFK, etc.)
Proficiency in at least one programming language (Python, Go, JavaScript, etc.)
Experience participating in on-call rotations and improving systems based on incidents
Strong communication and collaboration skills
You might be a great fit if you
Default to automation over manual processes
See repetitive work and immediately want to eliminate it
Think in terms of systems, failure modes, and long-term scalability
Care about building infrastructure that other engineers can use safely and confidently
Enjoy working in a small team with high ownership and impact
Nice to have
Experience running Kubernetes in production at scale
Deep familiarity with AWS
Experience building internal platforms or developer tooling
Background in distributed systems or large-scale data systems
We're a lean team, so your impact will be felt immediately, and opportunities will grow as the company scales up. If this all sounds like a good fit for you, why not join us?
Alloy is committed to fair and equitable compensation practices. Below is the anticipated starting base compensation range for this...