Design and build large-scale distributed services that improve fleet reliability, hardware health, and operational efficiency. Develop telemetry and analytics platforms that process and analyze infrastructure health signals at hyperscale. Build predictive models and intelligent services for hardware failure detection, repair recommendation, anomaly detection, and fleet risk forecasting. Analyze telemetry from servers, storage platforms, networking equipment, rack infrastructure, and datacenter systems to identify opportunities for improving reliability and availability. Partner with hardware, reliability, and capacity planning teams to develop data-driven operational strategies. Build AI-assisted experiences that accelerate incident investigation, root cause analysis, and repair decision-making. Design and implement safe automation and remediation workflows that reduce operational burden while maintaining strong operational controls. Participate in architecture reviews, code reviews, and live-site operations. Mentor engineers and contribute to engineering excellence across the organization. Drive projects from design through deployment and operational ownership. Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python These requirements include but are not limited to the following specialized security screenings: Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Experience designing and operating distributed systems and cloud services at scale. Experience working with hardware infrastructure, storage systems, server platforms, networking systems, or datacenter operations. Experience using data science, statistics, machine learning, forecasting, anomaly detection, or predictive analytics to solve engineering problems. Experience with telemetry and data platforms such as Azure Data Explorer (Kusto), Spark, Fabric, Databricks, or similar analytics technologies. Experience developing AI-powered operational tools, intelligent automation systems, or agent-based solutions. Experience with hardware reliability engineering, fleet management, capacity planning, or infrastructure health monitoring. Experience working with M365 components like Exchange, Substrate, SharePoint to improve performance, availability and supportability of services. Demonstrated ability to independently drive complex technical projects from concept through production deployment. Collaboration and communication skills with the ability to influence across organizations.
Want jobs like this matched to you?
SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.