Partners with appropriate stakeholders to determine user requirements for one or more complex scenarios. Provides technical leadership for the identification of dependencies and the development of design documents for a product, application, service, or platform. Leads by example and mentors others to produce extensible and maintainable code used across the company. Leverages deep subject-matter expertise of cross-product features with appropriate stakeholders (e.g., project managers) to lead multiple product's project plans, release plans, and work items. Holds accountability as a Designated Responsible Individual (DRI), mentoring engineers across products/solutions, working on-call to monitor system/product/service for degradation, downtime, or interruptions. Proactively seeks new knowledge and adapts to new trends, technical solutions, and patterns that will improve the availability, reliability, efficiency, observability, and performance of products while also driving consistency in monitoring and operations at scale and shares knowledge with other engineers. Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python These requirements include but are not limited to the following specialized security screenings: Master's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 15+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Distributed systems depth: Deep, hands-on expertise designing and shipping large-scale distributed systems, cloud services, or data platforms in production. Applied AI/ML: Demonstrated experience applying AI/ML — including modern LLM-based approaches — to real production problems, from prototype through delivery. Telemetry & big data: Proven ability to work with large-scale telemetry, observability, or big-data pipelines (metrics, logs, traces, time-series, anomaly detection). Delivery-led influence: A track record of leading through delivery and technical influence across teams and organizational boundaries, without relying on formal authority. AIOps & incident systems: Experience building AIOps, incident management, monitoring/detection, root-cause diagnosis, or automated mitigation systems for high-scale services. Telemetry for machines: Experience designing telemetry and instrumentation intended to be consumed by ML/LLM systems, not only by human operators. Agentic AI patterns: Familiarity with agentic AI patterns — planning, tool use, and autonomous or semi-autonomous action in production workflows. Foresight: A history of anticipating industry and technology shifts and repositioning an organization to capitalize on them ahead of peers.
Want jobs like this matched to you?
SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.