High Performance Compute Systems Site Lead (Onsite - LANL)

All, NMFull-timePosted Aug 5, 2026
High Performance Compute Systems Site Lead (Onsite - LANL)

  

This role has been designated as ‘Remote/Teleworker’, which means you will primarily work from home.

Who We Are:

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.

Job Description:

   

Join Hewlett Packard Enterprise as a Technical Site Lead supporting some of the world’s most advanced high-performance computing (HPC) and AI systems at Los Alamos National Laboratory (LANL), including large-scale Linux and HPE Cray environments. This senior, hands-on individual contributor position provides onsite technical and service-delivery leadership for HPE hardware engineers, Linux system administrators, software analysts, inventory specialists, and remote engineering resources supporting mission-critical HPE Cray and related platforms that enable scientific discovery and national security.

This is a technical leadership role with no direct reports. The Technical Site Lead combines hands-on Linux systems administration and HPC support experience with enterprise server hardware capability, structured troubleshooting, customer-facing leadership, incident coordination, and disciplined service delivery. The role is accountable for coordinating the onsite team's technical execution, maintaining operational readiness and service quality, managing escalations, planning maintenance, supporting major incidents, and preparing the site for next-generation HPC and AI platforms.

Operating as the site's technical leader, this individual establishes daily priorities, coordinates assignments, provides clear technical direction, mentors team members, follows through on commitments, and removes barriers affecting service delivery. Success requires the ability to lead through influence, work effectively across onsite and remote engineering organizations, communicate clearly with customer and HPE stakeholders, and maintain consistent execution in a complex, mission-critical computing environment.

US Citizenship and the ability to obtain and maintain a DOE Q Clearance are required.

Daily onsite work is required in Los Alamos, New Mexico. This is not a remote or hybrid position. Standard work hours are Monday through Friday, either 8:00 a.m. to 5:00 p.m. or 7:00 a.m. to 4:00 p.m., with additional onsite support required for planned maintenance, major incidents, and on-call lead responsibilities.

Key Responsibilities

Technical Leadership & Service Delivery

  • Provide day-to-day technical leadership and technical guidance to onsite HPE hardware engineers, Linux system administrators, and software analysts, while coordinating work across inventory specialists and remote engineering resources supporting large-scale HPE Cray EX and related HPC and AI systems.
  • Set and communicate daily and weekly technical priorities based on system health, open support cases, scheduled maintenance, customer priorities, operational risk, service commitments, and available staffing.
  • Serve as the primary onsite technical focal point for day-to-day system support, technical escalations, maintenance activities, and service-delivery risks, partnering with the DSM on customer governance, executive escalations, and broader service-delivery matters.
  • Maintain a current view of system health, support-case status, technical risks, maintenance actions, ownership, and outstanding commitments. Prepare and lead routine onsite operational reviews with the customer and onsite team in coordination with the DSM and HPC leadership.
  • Ensure support cases contain accurate technical details, diagnostic evidence, business impact, troubleshooting history, current ownership, and clearly defined next actions.
  • Drive timely escalation through established HPE processes and help team members engage next-tier support, engineering, product teams, and other resources needed to advance diagnosis and resolution.
  • Plan and coordinate system upgrades, maintenance windows, installations, acceptance activities, and other production-impacting work in accordance with customer change-control requirements and HPE support processes. Review prerequisites, risks, execution steps, and expected outcomes with the onsite team and customer before work begins.
  • Confirm required staffing, parts, tools, test equipment, documentation, communications, escalation contacts, rollback plans, and contingency coverage are in place for planned work.
  • Lead the HPE onsite response during major incidents by aligning technical priorities with the customer-designated incident lead and DSM, organizing HPE resources, maintaining clear communications, tracking actions and decisions, and ensuring appropriate escalation.
  • Coordinate HPE root-cause analysis, corrective actions, lessons learned, and documentation updates after significant incidents or recurring issues involving HPE-supported hardware, firmware, and infrastructure.
  • Identify risks to system availability, service-level commitments, maintenance schedules, operational readiness, or customer satisfaction and escalate them promptly to the DSM.
  • Coordinate onsite parts inventory, repair materials, tools, test equipment, and other HPE-owned resources using approved HPE business systems and controls.

Operational & Team Support

  • Build and maintain effective, professional working relationships with onsite team members, remote engineering organizations, HPE leadership, customer technical staff, and customer management.
  • Facilitate regular team coordination discussions to organize work, confirm ownership, review progress, surface technical blockers, and ensure commitments are completed or appropriately escalated.
  • Provide technical mentoring, coaching, and practical guidance to team members without assuming formal people-management authority.
  • Promote a culture of accountability, disciplined troubleshooting, accurate documentation, safe work practices, knowledge sharing, and professional customer engagement.
  • Track completion of required HPE, customer, security, safety, and technical training; notify team members of approaching deadlines and escalate overdue or at-risk requirements to the DSM.
  • Coordinate onsite coverage using approved team schedules and planned leave information. Identify and escalate potential gaps in business-hours, maintenance, or on-call coverage before they affect service delivery.
  • Provide the DSM with fact-based observations regarding technical performance, development needs, recognition opportunities, and issues requiring formal management attention.
  • Coordinate site-specific onboarding and operational readiness for new team members, including required accounts, badges, site access, training, workspace, equipment, and introductions to key stakeholders.
  • Maintain accurate site procedures, contact lists, escalation paths, team schedules, operational references, and other information required for effective day-to-day support.
  • Maintain a clean, safe, secure, and organized working environment in accordance with HPE and customer requirements.
  • Participate in the on-call rotation and provide additional onsite support when required for 24x7 operations, planned maintenance, system outages, and major incidents.

Hands-on Technical Contribution

  • Use Linux command-line and diagnostic tools in Red Hat Enterprise Linux (RHEL), SUSE Linux Enterprise Server (SLES), or comparable environments to analyze processes, filesystems, services, permissions, network state, system logs, hardware telemetry, and overall system health.
  • Lead and participate in monitoring, diagnosis, maintenance, and restoration of large-scale HPC compute, high-speed interconnect, storage, and management infrastructure, along with HPE-supported power and cooling components.
  • Apply a structured, evidence-based troubleshooting methodology that correlates Linux logs, hardware telemetry, network state, firmware information, and prior case history to isolate faults, test hypotheses, document findings, and determine the appropriate corrective action or escalation path.
  • Diagnose and support repair of enterprise server hardware, compute nodes, management components, HPE-supported interconnect and network components, storage components, power systems, cabling, and other integrated HPC equipment.
  • Perform or assist with hardware component replacement, cable and fiber management, rack-level work, equipment installation, electrostatic-discharge controls, and other hands-on data center activities in accordance with approved service procedures, safety requirements, and change controls.
  • Use out-of-band management controllers and interfaces, including BMCs, Redfish, and IPMI, to assess hardware health, validate firmware inventory, review environmental conditions and event logs, manage power state, and verify component status.
  • Read and interpret system documentation, hardware diagrams, rack elevations, network diagrams, cable maps, schematics, and service procedures to locate, identify, and diagnose system components.
  • Support new-system installation, expansion, integration, acceptance testing, hardware and firmware upgrades, and transition-to-operations activities.
  • Create and maintain site documentation, troubleshooting procedures, maintenance plans, workflows, technical checklists, incident records, and knowledge articles.
  • Use Bash, Python, Git, and other approved scripting, version-control, and collaboration tools to collect information, automate repeatable tasks, analyze system data, and maintain operational documentation and configuration references.

Required Qualifications

Candidates must meet the following core minimum requirements:

  • US Citizenship and the ability to obtain and maintain a DOE Q Clearance.
  • Must work onsite M-F in Los Alamos, New Mexico, with additional onsite work as required for planned maintenance, major incidents, and on-call support. This is not a remote or hybrid position.
  • High school diploma or equivalent with at least 7 years of relevant technical experience, or an associate or bachelor’s degree in a technical field with at least 5 years of relevant technical experience.
  • 5+ years of hands-on experience supporting complex electronic systems, enterprise server hardware, integrated data center infrastructure, or comparable production technology environments. This experience must include diagnosing, repairing, or maintaining enterprise server components, including processors, memory, storage devices, power supplies, BMCs, network adapters, optical connectivity, and copper or fiber cabling.
  • 3+ years of hands-on experience supporting HPC systems, supercomputing environments, large-scale Linux clusters, or similarly complex Linux-based compute infrastructure. This experience must include troubleshooting interactions among compute nodes, management systems, high-speed interconnects, storage platforms, operating systems, workload managers or schedulers, power, cooling, and supporting infrastructure.
  • 3+ years of experience providing technical leadership, mentoring, work coordination, or task direction for a multidisciplinary technical team. Direct people-management experience is not required.
  • 3+ years of hands-on Linux system administration, production support, or troubleshooting experience with Red Hat Enterprise Linux (RHEL), SUSE Linux Enterprise Server (SLES), or a comparable enterprise Linux distribution. Candidates must be able to independently use Linux command-line tools to navigate filesystems, inspect and manage processes and services, analyze system logs, review permissions, perform package-management tasks, validate network configuration and connectivity, and collect diagnostic information.
  • 2+ years of experience serving as a customer-facing technical focal point and coordinating support cases, maintenance activities, escalations, or incidents in a production environment using formal ticketing, change-control, incident-management, or service-management processes. Experience must include documenting impact and troubleshooting history, establishing ownership and next actions, coordinating onsite and remote resources, communicating status during major incidents or production outages, and following commitments through resolution.
  • Demonstrated experience using Bash, Python, or a comparable scripting language to collect system information, parse logs, automate repeatable operational tasks, analyze system data, or support troubleshooting activities.
  • Demonstrated hands-on experience safely using common hand tools, cable and fiber tools, electrostatic-discharge protection, and documented hardware service procedures to install, remove, inspect, cable, or replace rack-mounted server components.
  • Demonstrated ability to apply a structured, evidence-based troubleshooting process that includes defining the problem, collecting and preserving diagnostic evidence, isolating variables, testing hypotheses, documenting results, and determining the appropriate corrective action or escalation path.
  • Demonstrated ability to communicate technical information clearly to junior technical staff, experienced engineers, organizational leadership, customers, and remote support organizations. Candidates must have experience creating several of the following: technical procedures, maintenance plans, support-case updates, incident summaries, troubleshooting notes, knowledge articles, executive status summaries, and customer-facing communications
  • Experience supporting a 24x7 production environment and willingness to participate in an on-call rotation, planned after-hours maintenance, and additional onsite support during major incidents or system outages.
  • Demonstrated ability to read and interpret technical documentation, hardware diagrams, schematics, rack elevations, cable maps, network diagrams, and hardware service procedures to locate components, validate configurations, and support diagnostic or maintenance activities
  • Working proficiency with Windows or macOS and with standard browser-based, collaboration, and productivity tools, including Microsoft 365, SharePoint, Slack, Outlook, and Teams.
  • Ability, with or without reasonable accommodation, to work in datacenter and computer-room environments, lift up to 50 pounds independently and up to 75 pounds with assistance, perform rack-level and equipment-handling activities, and consistently follow safety, security, documentation, configuration-control, and information-protection requirements.

Preferred Qualifications

  • Prior DOE Q, DoD Top Secret, or comparable federal security-clearance experience. An active DOE Q Clearance is strongly preferred.
  • Hands-on experience supporting HPE Cray EX systems, HPE Cray Supercomputing platforms, or other leadership-class and exascale supercomputing environments.
  • Experience independently performing enterprise Linux diagnostics and administration involving system services, package management, network troubleshooting, log correlation, performance analysis, system-health assessment, and automation using Bash, Python, or comparable scripting tools.
  • Experience with HPE Slingshot interconnects, InfiniBand fabrics, high-speed Ethernet, or other large-scale HPC networking technologies.
  • Experience with high-speed network diagnostics, optical transceivers, copper and fiber cabling, link analysis, topology review, and fault isolation in large-scale environments.
  • Experience with liquid-cooled HPC infrastructure, cooling distribution units (CDUs), high-density compute cabinets, direct-liquid cooling, or related facility interfaces.
  • Experience using Redfish, IPMI, BMC interfaces, or comparable out-of-band management technologies for hardware monitoring, firmware inventory, event-log analysis, power control, and thermal review.
  • Experience supporting NVIDIA GB200 or GB300 NVL72 systems, NVIDIA DGX platforms, high-density GPU systems, accelerated computing platforms, or other rack-scale AI infrastructure.
  • Experience supporting parallel filesystems, enterprise storage platforms, cluster-management services, workload managers or schedulers, or HPC monitoring systems.
  • Project-management or technical workstream leadership experience coordinating upgrades, installations, maintenance windows, acceptance activities, or multi-party technical projects.
  • Experience with formal incident management, problem management, change management, service-level management, or ITIL-aligned service-delivery practices.
  • Experience using Git or a comparable version-control platform for scripts, configuration files, procedures, and technical documentation.
  • Experience working at a DOE national laboratory, DoD facility, government research organization, or another highly regulated customer environment.
  • Relevant industry certifications such as CompTIA Linux+, Server+, Network+, Security+, Red Hat Certified System Administrator (RHCSA), ITIL Foundation, or equivalent technical certifications.
     

This role offers the opportunity to provide critical onsite technical leadership for mission-critical HPC operations at one of the nation's leading research laboratories. The successful candidate will be a technically credible, hands-on leader who can guide experienced professionals, work effectively with demanding customers, maintain disciplined service delivery, and help prepare the site for its next generation of HPC and AI platforms.

What We Can Offer You:

Health & Wellbeing

We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.

Personal & Professional Development

We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have — whether you want to become a knowledge expert in your field or apply your skills to another division.

Unconditional Inclusion

We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.

Let's Stay Connected:

Follow @HPECareers on Instagram to see the latest on people, culture and tech at HPE.

#unitedstates

Job:

Services

Job Level:

TCP_04

    

"The expected salary/wage range for this position is provided below. Actual offer may vary from this range based upon geographic location, work experience, education/training, and/or skill level.
– United States of America: Annual Salary USD 105,500 - 243,000 in New Mexico
The listed salary range reflects base salary. Variable incentives may also be offered."

Information about employee benefits offered in the US can be found at https://myhperewards.com/main/new-hire-enrollment.html

HPE is an Equal Employment Opportunity/ Veterans/Disabled/LGBT employer. We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need. Our goal is to be one global team that is representative of our customers, in an inclusive environment where we can continue to innovate and grow together. Please click here: Equal Employment Opportunity.

Hewlett Packard Enterprise is EEO Protected Veteran/ Individual with Disabilities.

   

HPE will comply with all applicable laws related to employer use of arrest and conviction records, including laws requiring employers to consider for employment qualified applicants with criminal histories.

   

Recruitment Fraud Alert

We have become aware of an increase in fraudulent recruitment activities in which individuals impersonate our company or authorized recruitment agencies to offer fake employment opportunities. These scams may occur through false websites, emails, social media, or chat-based applications and often aim to obtain personal information or money. Please note that Hewlett Packard Enterprise (HPE), its direct and indirect subsidiaries and affiliated companies, and its authorized recruitment agencies/vendors will never charge a candidate a registration fee, hiring fee, or any other fee in connection with its recruitment and hiring process. We also never request personal information such as back account details, Social Security numbers, or national IDs via social media or chat applications.

All legitimate job opportunities will come through official company channels, and candidates are responsible for verifying the credentials of any third party claiming to represent the company. Any reliance on fraudulent communication is at the individual’s own risk, and HPE disclaims legal liability for any resulting damages. If you suspect recruitment fraud, do not share personal information or make any payments and report the incident to your local authorities immediately.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free