Build and maintain software components for monitoring, alerting, dashboards, diagnostics, telemetry collection, and service health validation. Implement automation for alert enrichment, incident context generation, health checks, and operational reporting. Use logs, metrics, traces, telemetry, tests, and debugging tools to investigate issues and improve diagnosability. Participate in code reviews and apply team standards for maintainability, reliability, testability, security, and privacy. Contribute to AI-assisted workflows for incident classification, evidence collection, root-cause investigation, and remediation recommendations. Participate in on-call rotations, incident retrospectives, and follow-up work that improves monitoring, tests, documentation, and troubleshooting guides. Escalate blockers, design issues, reliability risks, and production concerns with clear impact and timeline Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Software engineering experience building features, services, tools, or automation using languages such as C#, Python, Go, Java, PowerShell, or TypeScript. Foundational understanding of distributed systems, cloud services, APIs, monitoring, logging, telemetry, and production debugging. Ability to write maintainable, testable, diagnosable, and secure code with guidance from experienced engineers. Collaboration and communication skills in engineering and live-site situations. Experience or interest in Azure Monitor, Log Analytics, Application Insights, Kusto/KQL, Geneva, IcM, or similar observability and incident-management systems.
Want jobs like this matched to you?
SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.