Senior Software Engineer

United StatesPosted Aug 5, 2026

Define architecture for agentic reliability systems spanning monitoring, telemetry, incident management, service topology, deployment signals, and operational knowledge. Lead platform capabilities for automated detection, triage, root-cause assistance, mitigation recommendations, safe execution, and post-incident learning. Establish engineering standards for safe agentic operations, including identity, access, compliance, rollback, auditability, change management, and human escalation. Influence service teams to adopt consistent monitoring, SLOs, alert quality, incident automation, live-site readiness, and operational excellence practices. Identify high-impact reliability gaps and convert them into platform investments, architectural improvements, and reusable automation. Lead complex live-site investigations and drive systemic reliability improvements from incident patterns and customer-impact data. Mentor engineers and shape long-term technical direction across monitoring, observability, incident response, and AI-assisted and agentic automation. Embody our culture and values Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Experience building production-scale platforms, cloud services, distributed systems, or reliability automation. Experience architecting complex systems across service boundaries and driving execution across partner teams without direct authority. Experience with observability architecture, monitoring systems, incident response, service health modeling, operational automation, and production debugging. Solid judgment around production safety, automation risk, customer impact, security, privacy, compliance, and responsible AI-assisted and agentic automation. Experience leading agentic automation, AI-assisted diagnostics, autonomous remediation, intelligent operations, or reliability platform efforts.Experience with Azure Monitor, Log Analytics, Application Insights, Kusto/KQL, Azure Resource Graph, Azure DevOps, GitHub, or similar monitoring and observability ecosystems. Experience creating organization-level reliability metrics such as SLO compliance, alert quality, time to detect, time to mitigate, human effort saved, automation coverage, and incident recurrence. Experience mentoring senior engineers and defining technical strategy across monitoring, observability, incident response, and production engineering disciplines.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free