Skip to main content.
sitemap
Profile
Sign Out
View More Jobs
Manager, Site Reliability Engineering
Reston, VA, United StatesAustin, TX, United States
Be the First to Apply
Job Identification
340413
Job Category
Product and Research
Posting Date
07/23/2026, 05:58 PM
Role
People Manager
Job Type
Regular Employee
Does this position require a security clearance?
Yes
Years
6 to 10+ years
Applicants
Less than 10 applicants
Additional Info
Visa / work permit sponsorship is not available for this position
Applicants are required to read, write, and speak the following languages
English
Job Description
Supports team members in designing and architecting infrastructure and service for reliability and functionality. Provides day-to-day direction to help forecast demands and ensure systems have adequate resources. Assists with collaboration between team members and the software development team to create reliable, scalable infrastructures. Advises on data collection, optimizing operations and infrastructure reliability. Aids in incident response activities to ensure service reliability. Reviews health and performance reports. Trains team members to identify automation. Trains team members to communicate and understand the impact of changes. Serves as an escalation point for incidents and reviews documentation. Enables team members to experiment with new technology, execute improvements, build site reliability knowledge, and provide clear data.
Responsibilities
Key Responsibilities
Capacity Ingestion and
Management:
-
Supports
team members designing and architecting infrastructure and/or service according
to terms for reliability and functionality.
-
Supervises
immediate team members and provides day-to-day direction to help forecast
demands for infrastructure and respond to capacity needs, ensuring systems have
sufficient resources to handle current and future workloads.
-
Assists
team members in collaborating with the software development team to develop
infrastructures and features that are reliable and scalable according to
deployment requirements.
-
Develops
team members' ability to identify opportunities for and drive prototyping
(e.g., testing new applications or infrastructures, assisting in onboarding).
Incident and Service
Lifecycle Management:
-
Advises
team members on performing data collection, triage, technical analysis, and
redirection and recommends methods to maintain and optimize operations and
infrastructure reliability.
-
Provides
support to team members monitoring services, ensuring they maintain up-to-date
knowledge of performance and document their condition.
-
Leverages
working knowledge to aid team members in performing incident response, root
cause analyses, and/or maintenance on assigned services (e.g., software
installs, version upgrades, security updates, backup and recovery).
-
Reviews
health and performance reporting and recommends appropriate actions based on
trends in data.
-
Helps
team members follow procedures to perform provisioning to support
infrastructure, applications, and services.
-
Educates
team members on performing decommissioning (e.g., shutting down servers,
removing data from databases) to remove objects that are no longer needed.
Automation:
-
Trains
team to identify opportunities for automation and assesses potential benefits.
-
Reviews
automation tools or scripts developed by team...
Only part of this posting is shown here. Read the full description