Site Reliability Engineer, Litmus GCC
Who is Litmus
Litmus is building the data foundation that powers industrial AI.
AI doesn’t work without real-world, contextualized data - Litmus makes that data usable. As AI adoption accelerates, most industrial environments still can’t access or use their operational data. We solve that gap.
We’re a growth-stage software company helping manufacturers access, structure, and use real-time data from machines, systems, and sensors at the edge. Our platform sits at the intersection of edge computing, AI, and industrial operations, enabling some of the world’s largest companies to run operations in real time, reduce downtime, and optimize production.
Backed by leading investors and trusted by global manufacturers and partners like Google, Microsoft, Dell, Oracle, and Mitsubishi, Litmus is powering the shift toward software-defined manufacturing.
Why join Litmus
Build the infrastructure that makes industrial AI possible
AI is moving beyond the cloud and into the physical world. At Litmus, you’ll build the infrastructure that enables real-time data to power AI and machine learning systems in production environments.
Work on problems where software meets the real world
Most AI systems fail without access to real-world data. You’ll build the layer that makes them viable in production. We solve challenges at the intersection of distributed systems, real-time data, and industrial constraints — where reliability, scale, and performance are non-negotiable.
Have real impact, fast
You’ll work on systems used by real customers in production, with direct impact on product and company trajectory. As a scaling company, we move quickly. You’ll have ownership, visibility, and the ability to shape both product and company as we scale.
Join a high-performance team
We’re building a team that holds a high bar and pushes each other to improve. You’ll work alongside experienced operators, engineers, and leaders who have done this before and are building again at scale. We hire people who take ownership, move quickly, and care about outcomes. No passengers.
Our culture
At Litmus, the team is collaborative, curious, and low ego. People are scrappy, take ownership, and look for ways to make an impact. We value empathy just as much as execution, whether that’s in how we build, how we communicate, or how we support each other.
We’re a growing company, so things move quickly and not everything is perfectly defined. If you enjoy figuring things out, working closely with others, and making steady progress, you’ll do well here.
About the Role
Litmus Automation is hiring a Site Reliability Engineer to own the day-to-day reliability, security, and performance of our Azure-hosted Litmus Unified Namespace (UNS) and Litmus Edge Manager (LEM) environment, deployed for a strategic enterprise customer. This role covers the full lifecycle: initial environment provisioning and hardening, deployment support, ongoing monitoring and incident response, and long-term operation against a 99.9% uptime commitment. You will work directly with customer-facing infrastructure spanning AKS, Azure networking, managed databases, identity federation, and MQTT-based data pipelines connecting multiple sites. As this is an SRE role held to a 99.9% uptime commitment, it includes participation in an on-call rotation to ensure continuous coverage and rapid incident response outside of standard working hours.
Key Responsibilities
• Provision, configure, and maintain cloud infrastructure spanning compute, container orchestration, networking, and database services.
• Own end-to-end monitoring and alerting, and build/maintain operational dashboards that give clear visibility into system health and performance.
• Drive response and on-call support to meet uptime SLAs, including root cause analysis and post-incident reviews.
• Implement and maintain a strong security baseline, including vulnerability management, secrets management, network security controls, and certificate lifecycle management.
• Configure and support identity/SSO integration between internal and customer identity providers, troubleshooting authentication issues as needed.
• Manage networking and connectivity to support reliable, high-throughput data flows across distributed environments and sites.
• Monitor and report on infrastructure usage and consumption against agreed baselines; support capacity planning.
• Build and maintain infrastructure-as-code and CI/CD pipelines to automate provisioning and deployment.
• Manage backup and disaster recovery processes and periodically validate recovery procedures.
• Participate in customer-facing environment validation, UAT support, and go-live sign-off activities.
• Maintain clear runbooks, documentation, and knowledge transfer materials for the environment.
Required Qualifications
• 4-8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, with significant hands-on Azure experience.
• Strong hands-on experience with core Azure services: AKS, Virtual Machines, VNet/NSG/Load Balancer/Application Gateway, Azure Database for PostgreSQL/MySQL, Key Vault, Azure Monitor, and Log Analytics.
• Practical experience administering and troubleshooting Kubernetes’ workloads in production.
• Experience with infrastructure-as-code tools such as Terraform, Bicep, or ARM templates.
• Proficiency in scripting/automation using Python, Bash, or PowerShell.
• Solid understanding of networking fundamentals (DNS, TLS/SSL, load balancing, firewalls/NSGs).
• Experience with incident management, on-call rotations, and SLA-driven operations.
• Strong communication skills, with the ability to work directly with customer stakeholders during validation and escalations.
Preferred / Nice-to-Have
◦ Microsoft Certified: DevOps Engineer Expert (Azure DevOps certification) is required.
◦ Experience with identity federation and SSO protocols (OIDC/SAML), particularly Key cloak and/or Okta.
◦ Familiarity with MQTT or other IoT/industrial data protocols.
◦ Additional certifications: Microsoft Certified: Azure Administrator Associate, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).
◦ Experience supporting manufacturing, industrial, or IoT customer environments.