Site Reliability Engineer (SRE)
Remote$145k–$160kPosted Jul 18, 2026
All Jobs
>
Site Reliability Engineer (SRE)
Apply
Site Reliability Engineer (SRE)
Remote - US
Apply
Description
About PaylianceFounded in 2007, Payliance is a trusted leader in payment processing — processing more than $63 billion annually, supporting 40,000+ merchant locations, and serving over 350 lending clients. We offer an all-in-one platform for real-time funding, payment processing, account verification, and recovery services, giving lenders the technology to operate efficiently and confidently.What sets Payliance apart is our blend of modern technology, deep industry expertise, and a highly collaborative, people-first culture. Backed by Serent Capital, we're expanding our capabilities and delivering measurable results for clients across lending, e-commerce, collections, and gaming.About the RoleThe Site Reliability Engineer (SRE) bridges software engineering and infrastructure operations, owning the reliability, scalability, and performance of Payliance's payment processing platform. You'll apply engineering discipline to operational problems — reducing toil, automating repetitive work, and building systems that are resilient by design.This role is ideal for an engineer who can read and debug production .NET code, has hands-on AWS compute and RDS SQL Server depth, and brings the observability mindset needed to keep a high-transaction payment platform running at scale.What You'll Do.NET Application Reliability· Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues.· Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents.· Partner with application engineers to embed reliability into new feature design and deployment practices.· Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early.AWS Compute & Infrastructure· Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads — with clear judgment on when each is the right fit.· Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments.· Automate operational tasks, deployment pipelines, and disaster recovery procedures.· Continuously reduce toil through tooling and automation, freeing the team for higher-impact engineering work.RDS SQL Server Operations· Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup.· Operate backup and point-in-time recovery (PITR) processes and validate restore procedures regularly.· Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains.· Capacity plan and scale database infrastructure to support transaction volume growth.Observability & Monitoring· Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms; AWS X-Ray for distributed tracing.· Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements.· Design alerts that surface signal — not noise — and ensure on-call responders have the context to act quickly.· Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence.Networking & Security· Design and maintain secure AWS network topologies: VPCs, subnets, security groups, and NACLs.· Configure and manage ALB/NLB routing, Route 53 DNS, and TLS certificate lifecycle via ACM.· Author and review least-privilege IAM policies; audit roles and resource-based policies for over-permissioning.· Support compliance and security controls relevant to a PCI-regulated payments environment.Incident Response & On-Call· Participate in on-call rotation to respond to production incidents and drive swift resolution.· Define and track error budgets; use them to balance velocity and reliability investment.· Communicate status updates clearly during incidents and coordinate...