Senior Reliability Developer
At Arctic Wolf, you won’t just watch the cybersecurity industry evolve – you'll help lead the change. Our global Pack is made up of people who thrive on solving hard problems, moving fast, and building technology that protects organizations around the world. We’re proud to be recognized by Forbes, CNBC, Fortune, CRN, Gartner Peer Insights and IDC MarketScape – but what matters most is the work behind it: delivering real outcomes for customers through award winning innovation like our Aurora Platform.
If you’re looking for meaningful work, smart teammates and the chance to make a real impact in a high-growth company that’s redefining security operations, Arctic Wolf is the right place for you!
Our mission is simple: End Cyber Risk. We’re looking for a Senior Reliability Developer to be part of making this happen.
LOCATION: 8939 Columbine Road, Eden Prairie, MN 55347; Telecommuting permissible from any location in the US.
About the Role
Responsible for changes and updates to public/private cloud infrastructure. Owns all monitors that support AWN business services and our customers. Implements effective automate using programming languages such as Bash, Python, and JavaScript. Ensure our full infrastructure stack is resilient with day-to-day care and feeding. Write and manage Terraform modules and configuration trees for deployments into AWS and OpenStack. Manage automated patching solutions for instances deployed across multiple regions to ensure compliance with strict security requirements. Monitor internal and external TLS certificates for expiration, renewals, and re-deployment into their respective environments. Create and improve service monitoring solutions using Prometheus, Grafana, Zabbix, and various Prometheus exporters (e.g., postgres, node-exporter). Migrate monitoring for legacy production services to a new monitoring system and recreated existing monitoring within Prometheus + Grafana. Collaborate frequently with development teams to implement new features, fixes, or troubleshoot services in lab/production environments. Troubleshoot problems with microservices running within containers, administer containers in Kubernetes clusters, manage cloud components deployed in multiple AWS regions, and maintain CI/CD pipelines. Manage the Cylance AWS Amazon cloud platform through regular administration of Kubernetes cluster upgrades, cluster module upgrades, logging configurations, and resizing of EBS volumes and instances. Troubleshoot issues between interconnected services by working with various teams and resolving issues related to API calls and service unavailability. Create, improve, and maintain detailed documentation for MOPs, service architecture, runbooks, monitoring configurations, and other general resources. Plan and execute service decommissioning in both lab and production environments upon customer or service owner request. Work on an on-call rotation to support business-critical services outside of business hours, respond to escalations, and conduct root cause analysis during incident reviews. Telecommuting permissible from any location within US.
Requirements: Bachelor’s degree or foreign degree equivalent in Computer Information Systems, or related field and five (5) years of progressive, postbaccalaureate experience in a Technology related role or job offered or related role.
Experience and/or education must include:
Utilizing Infrastructure-as-Code (IaC) tools including Terraform and configuration management tools including Puppet and Chef to design, implement, and manage multi-region cloud infrastructure deployments across AWS and OpenStack environments.
Utilizing containerization and orchestration technologies including Docker, Kubernetes, ECS, and EKS to deploy, manage, and troubleshoot microservices architectures in production environments.
Utilizing Python, Bash, JavaScript, Groovy, and PowerShell scripting languages for automation of infrastructure provisioning, certificate lifecycle management, and automated patching solutions.
Utilizing monitoring and observability tools including Prometheus, Grafana, Zabbix, AlertManager, and PagerDuty to create dashboards, configure alerts, and implement service health monitoring solutions.
Utilizing AWS cloud services including EC2, ECS, EKS, ELB, S3, RDS, IAM, Lambda, CloudFormation, VPC, Route53, and CloudWatch to architect, deploy, and manage highly available cloud-based systems.
Utilizing CI/CD pipeline tools including Jenkins, GitLab, GitHub, Bitbucket, and Gaia to automate testing, deployment processes, and maintain continuous integration workflows.
Managing Kubernetes cluster operations including cluster upgrades, module upgrades, logging configurations, node scaling, EBS volume management, and pod orchestration across multiple AWS regions.
Utilizing certificate management and security tools including Keeper for secret management, TLS certificate monitoring, renewal coordination, and automated deployment across production and development environments.
Troubleshooting distributed systems and microservices communication issues including API failures, service mesh architectures, container networking, inter-service dependencies, and deployment failures.
Creating and maintaining technical documentation using Atlassian tools including Jira, Confluence, Backstage, and LucidChart for MOPs (Method of Procedures), runbooks, service architecture diagrams, and monitoring configurations.
About Arctic Wolf
At Arctic Wolf, we foster a collaborative and inclusive work environment that thrives on diversity of thought, background, and culture. This is reflected in our multiple awards, including Top Workplace USA (2021-2025), Best Places to Work – USA (2021-2025), Great Place to Work – Canada (2021-2024), Great Place to Work – UK (2024-2026), and Kununu Top Company – Germany (2024-2026). Our commitment to bold growth and shaping the future of security operations is matched by our dedication to customer satisfaction, with over 10,000 customers worldwide and more than 2,000 channel partners globally. As we continue to expand globally and enhance our technology, Arctic Wolf remains the most trusted name in the industry.
Our Values
Arctic Wolf recognizes that success comes from delighting our customers, so we work together to ensure that happens every day. We believe in diversity and inclusion, and truly value the unique qualities and unique perspectives all employees bring to the organization. And we appreciate that—by protecting people’s and organizations’ sensitive data and seeking to end cyber risk— we get to work in an industry that is fundamental to the greater good.
We celebrate unique perspectives by creating a platform for all voices to be heard through our Pack Unity program. We encourage all employees to join or create a new alliance. See more about our Pack Unity here.
We also believe and practice corporate responsibility, and have recently joined the Pledge 1% Movement, ensuring that we continue to give back to our community. We know that through our mission to End Cyber Risk we will continue to engage and give back to our communities.
All wolves receive compelling compensation and benefits packages, including:
Equity for all employees
Flexible time off and paid volunteer days
RRSP and 401k match
Training and career development programs
Comprehensive private benefits plan including medical, mental health, dental, disability, life and AD&D, and value-added services
Robust Employee Assistance Program (EAP) with mental health services
Fertility support and paid parental leave
Arctic Wolf is an Equal Opportunity Employer and considers applicants for employment without regard to race, color, religion, sex, orientation, national origin, age, disability, genetics, or any other basis forbidden under federal, provincial, or local law. Arctic Wolf is committed to fostering a welcoming, accessible, respectful, and inclusive environment ensuring equal access and participation for people with disabilities. As such, we strive to make our entire experience as accessible as possible and provide accommodations as required for candidates and employees with disabilities and/or other specific needs where possible. Please let us know if you require any accommodation by emailing recruiting@arcticwolf.com. View our Hiring Page to learn more about our application process.
Security Requirements
Conducts duties and responsibilities in accordance with AWN’s Information Security policies, standards, processes, and controls to protect the confidentiality, integrity and availability of AWN business information (in accordance with our employee handbook and corporate policies).
Background checks are required for this position.
This position may require access to information protected under U.S. export control laws and regulations, including the Export Administration Regulations (“EAR”). Please note that, if applicable, an offer for employment will be conditioned on authorization to receive software or technology controlled under these U.S. export control laws and regulations.
The base salary range for this job family is 145,018 to 180,000 USD annually. This range reflects the base pay the company reasonably expects to offer for this position, aligned to the broader job family base pay structure. Actual base pay may vary based on skills, experience, and location, including job family level. In addition to base pay, Arctic Wolf offers variable incentive compensation, new hire equity grants, and a comprehensive benefits package.