Principal Software Engineer - HW/SW in Fleet Infrastructure

United StatesPosted Jul 27, 2026

Partner with broad teams to design, develop, validate, debug, and optimize infrastructure software across OS, firmware, fleet health, hardware repair, telemetry, and monitoring areas. Lead hardware and firmware reliability architecture, including health telemetry, diagnostics, failure detection, predictive insights, and remediation automation. Drive failure analysis and root cause investigation for critical live-site incidents involving hardware, firmware, OS, storage, networking, or infrastructure platforms, and drive durable improvements that reduce recurrence. Establish engineering standards, observability patterns, and operational mechanisms that improve fleet availability, reduce deployment risk, and strengthen hyperscale infrastructure operations. Evolve M365 substrate infrastructure strategy for AI and agentic workloads across compute, memory, storage, and networking domains. Influence cross-organizational technical direction, communicate complex tradeoffs clearly, and help teams make durable architecture decisions. Mentor early in career engineers. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python These requirements include but are not limited to the following specialized security screenings: Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Experience leading architecture and technical strategy for large-scale distributed systems, cloud infrastructure, or hyperscale service platforms. Deep experience with hardware platforms, firmware, OS, drivers, datacenter infrastructure software, storage systems, networking, or platform software. Experience with New Product Introduction, platform validation, deployment readiness, lifecycle management, or fleet-scale hardware operations. Experience on hyperscale fleet live site, using telemetry, observability, failure analysis, predictive diagnostics, or automation to improve infrastructure reliability. Experience working with silicon providers, hardware vendors, OEMs, ODMs, firmware teams, or platform engineering organizations. Experience influencing cross-organizational engineering strategy and driving complex technical programs across multiple teams. Communication skills with the ability to explain technical tradeoffs to senior engineering and business leaders.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free