Principal, System Reliability Engineer

Ayar Labs
San Jose, CA$185k–$290kPosted Jul 29, 2026
Position:  Principal, System Reliability Engineer Location:  San Jose, CA Job Id:  650 # of Openings:  0 Principle Systems Reliability Engineer   Location: San Jose , CA (On-site)

Ayar Labs is shattering AI data bottlenecks by moving data at the speed of light. As pioneers of co-packaged optics (CPO), we are using light instead of electricity to move data faster, further, and with a fraction of the energy needed to fuel the explosive growth of AI models.

Backed by industry giants like NVIDIA, AMD, Mediatek and Intel and manufactured in partnership with the world’s leading semiconductor ecosystem, Ayar Labs’ co-packaged optics solution is key to unleashing next-generation AI scale-up architectures. 

About the Role
Ayar Labs builds optical I/O technology for hyperscale AI infrastructure. This role owns reliability qualification at the system and fleet level: building the test infrastructure and evidence base that proves our products are ready for large-scale deployment, and then partnering directly with customers and integration partners to carry that evidence through joint qualification programs and pilot deployments. You will be responsible for both the engineering (test infrastructure, statistical reliability planning, standards compliance) and the relationship (translating technical evidence into the specific commitments our partners need to move forward).

What You'll Own
  • Test infrastructure and evidence generation: Design, build, and operate custom test infrastructure that emulates real system-level electrical, thermal, and workload conditions ahead of, or independent of, any single customer's specific hardware. This includes sourcing or generating representative workload and power profiles and using them to drive continuous, long-duration test campaigns.
  • Reliability statistics and demonstration planning: Define statistical sample sizes and test durations needed to demonstrate specific reliability and confidence targets, and apply the right methodology to each failure population in the system rather than a single blanket target.
  • Standards and compliance: Ensure that test infrastructure, interfaces, and telemetry are built against relevant industry interconnect and management standards, so evidence generated internally holds up when reviewed by, or transferred to, an external partner's platform.
  • Staged qualification execution: Execute staged qualification gates, from lab-level interoperability and functional testing through environmental and stress testing to live pilot deployment, applying a reliability / availability / serviceability lens throughout rather than treating any one dimension in isolation.
  • Fault injection and lifecycle monitoring: Design and run fault-injection and failure-mode testing, including scenarios that exercise field-serviceable components under representative operating conditions, in partnership with customer and integration-partner operations teams.
  • Economic modeling and fleet reporting: Provide reliability and performance data into cost-of-ownership and total-cost modeling shared with customers, and own ongoing fleet reliability reporting and incident response once products reach production.
  • Cross-functional partnership: Serve as the primary technical point of contact with Tier-1 integration partners and hyperscale customer engineering teams across the qualification lifecycle, from early technical evidence through pilot sign-off and into steady-state operations.

Basic Qualifications
  • 5+ years in fleet/systems reliability engineering for large-scale infrastructure, with a BS in Electrical Engineering, Computer Engineering, or a related field.
  • You've built or operated test infrastructure that validates a component or subsystem's behavior before it is deployed into a customer's actual system, using representative rather than production hardware.
  • You've defined and executed statistical reliability demonstration plans (sample sizes, test durations, confidence and reliability targets) for hardware components, and can explain the methodology behind them, not just apply a lookup table.
  • You've worked with die-to-die, chip-to-chip, or board-level interconnect standards and telemetry or management interfaces relevant to high-speed data center hardware.
  • You've partnered directly with external OEM or hyperscale customer engineering teams on a joint qualification or certification program, and are comfortable owning that relationship technically.
  • Comfort operating a live, continuously running test or fleet environment: on-call coverage, incident response, and building the telemetry pipeline that turns raw sensor data into fleet-level reliability statistics.
  • Strong written and verbal communication skills; you will regularly translate internal engineering data into evidence and documentation for external partners.

Preferred Qualifications
  • MS in Electrical Engineering, Computer Engineering, or a related field.
  • Experience with FPGA-based test or signal-generation systems.
  • Experience characterizing or emulating real compute or network workload behavior for use in a test or validation environment.
  • Background in data-center fleet operations, site-reliability engineering, or customer-facing qualification and certification processes.
  • Familiarity with co-packaged optics, silicon photonics, or other emerging interconnect packaging technologies.
  Salary Range:  $185,000 - $290,000     NOTE TO RECRUITERS:
Principals only. We are not accepting resumes from recruiters for this position. Remuneration for recruiting activities is only applicable subject to a signed and executed agreement between the parties. Please don’t send candidates to Ayar Labs, and do not contact our managers.

Ayar Labs is an Equal Opportunity Employer and is strongly committed to all policies which will afford equal opportunity employment to all qualified persons without regard to age, sex, national origin, race, color, ethnicity, creed, religion, gender identity, sexual orientation, disability, veteran status, or any other characteristic protected by law. It is the policy of Ayar Labs to provide reasonable accommodation when requested by a qualified applicant or employee with a disability, unless such accommodation would cause an undue hardship. Veterans are more than welcome and encouraged to apply.


Apply for this Position

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free