Develop telemetry, analytics, and reporting platforms that measure fleet compliance, update readiness, vulnerability exposure, firmware health, and remediation progress. Build intelligent services that detect freshness gaps, identify at-risk infrastructure, prioritize remediation, and forecast fleet security or compliance risk. Create safe automation workflows for operating system updates, firmware upgrades, security patch deployment, configuration enforcement, and remediation orchestration. Partner with security, hardware, firmware, operating system, and infrastructure teams to define and execute data-driven strategies for improving fleet health and reducing exposure. Build AI-assisted experiences that accelerate investigation of patching failures, firmware issues, configuration drift, vulnerability exposure, and live-site incidents. Improve operational safety through staged rollout systems, health gates, rollback mechanisms, policy enforcement, and strong observability. Participate in architecture reviews, code reviews, design discussions, and live-site operations. Mentor engineers and contribute to engineering excellence across the organization. Drive projects from concept and design through production deployment, measurement, and operational ownership. Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python These requirements include but are not limited to the following specialized security screenings: Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. In-depth knowledge of operating system internals, with strong expertise in Windows and practical experience with Linux, including debugging and troubleshooting kernels, drivers, system services, networking stacks, crash dumps, and performance issues. Experience collecting, correlating, and analyzing diagnostic data at scale, including event logs, traces, performance counters, networking telemetry, driver diagnostics, and other system-level signals to identify root cause and drive remediation. Proven ability to build or leverage pipelines that aggregate and analyze this data (e.g., Kusto/Azure Data Explorer, or similar) to identify root cause, detect systemic issues, and drive remediation across the M365 Fleet. Demonstrated ability to diagnose complex issues involving OS crash dumps, driver failures, networking (TCP/IP, DNS, packet capture), and system performance (CPU, memory, I/O, latency) using tools such as WinDbg, Windows Performance Analyzer, ETW (Event Tracing for Windows). Experience with security patching, vulnerability management, compliance reporting, configuration management, or secure infrastructure operations. Experience developing automation systems for safe deployment, remediation, rollout orchestration, rollback, or fleet-wide policy enforcement. Experience applying data science, statistics, forecasting, anomaly detection, or machine learning to infrastructure health, security posture, or operational risk problems. Experience developing AI-powered operational tools, investigation assistants, or agent-based solutions. Demonstrated ability to independently drive complex technical projects from concept through production deployment. Collaboration and communication skills with the ability to influence across teams and organizations.
Want jobs like this matched to you?
SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.