Principal Software Engineer

Microsoft
United StatesPosted Jul 29, 2026

Full-lifecycle engineering ownership: strength on both sides of the code. That means the up-front work (design, failure-mode and threat analysis, testability and operability designed in from the start) and everything after it lands (validation, safe rollout, monitoring, on-call support, and long-term sustainment). Writing the feature is the middle third of the job, not the whole of it. Production excellence / SRE depth: demonstrated ownership of a high-scale live service, including on-call leadership and incident ownership during high severity incidents, service health engineering (SLOs/SLIs, monitoring and alerting design, telemetry-driven diagnostics), safe deployment practices, and a track record of driving root cause through to permanent fixes rather than mitigations. Test engineering depth: designs test strategy for distributed systems end to end, covering integration, fault injection/chaos, and performance/load validation, and builds the test infrastructure and CI/CD quality gates behind it (deterministic suites, environment and data management, flakiness elimination). Treats correctness and testability as design-time concerns rather than a pre-release activity. As a Principal Software Engineer on the team, you will have the opportunity to work with various Azure technologies to build a massively scalable cloud service. You will develop and validate various components needed to build a robust, distributed and resilient platform for Azure Usage Billing along with helping in the design of parts of the platform. This includes working on service management, programmability, usage pipeline, service fundamentals like monitoring, security, performance, engineering systems, tooling and live site. Embody our culture and values Bachelor's Degree in Computer Science OR related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. 4+ years technical engineering experience with C# and the .Net Framework and associated ecosystem. Proven Site Reliability Engineering (SRE) experience operating mission-critical cloud services, with ownership of on-call excellence, incident management, observability, reliability improvements, and safe service deployment practices. Experience running and testing large scale platforms in a cloud across multiple regions and clouds. Experience in using ML/AI models to support and accelerate all parts of the service lifecycle from planning to support and sustainment. Experience working with data plane infrastructure and SDK creation across one or more languages. Experience building complex distributed systems and sync/async message based platforms. An understanding of cloud development principles and patterns. Communications skills and ability to work collaboratively across several teams. Problem-solving skills with ability to quickly adapt to new technology and go deep.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free