Principal Software Engineer
United States · California · Washington · Mountain View, CA · Redmond, WA · New York, NY$143k–$275kPosted Jul 25, 2026
Principal Software Engineer | Microsoft Careers
Skip to main content
Microsoft
Careers
Careers
Careers
Search Search jobs
Sign in
Single PositionView All JobsPrincipal Software EngineerUnited States, Washington, Redmond +2 moreApply nowAdd to cartFind out how well you match with this jobUpload your resumeJob descriptionCompany and benefitsJob number200044430Date postedJul 24, 2026Work site3 days / week in-officeTravelLess than 25%ProfessionSoftware EngineeringDisciplineSoftware EngineeringRole typeIndividual ContributorEmployment typeFull-TimeOverviewHow do you know an AI product is actually getting better, and how do you prove it, at scale, before millions of people feel the difference? We build and operate the offline evaluation platform that gates Microsoft Copilot’s quality: teams across Copilot depend on us to run their scenarios against the product, score the responses, and produce the scorecards that decide what ships. We do this at the scale of one of the world’s largest AI products, inside a strict enterprise compliance boundary. As Copilot becomes a fleet of autonomous agents, the way we measure quality has to be reinvented, and you will lead that reinvention. As a Principal Software Engineer on the Copilot evaluation team, you will own the technical vision for a platform built AI-first and agentic-first: one where autonomous agents are not just measured, but help operate, extend, and heal the system itself. You will take on our hardest problem, delivering trustworthy results across a deep chain of dependencies (Copilot’s model, retrieval, and scoring services) that we drive far outside their normal operating envelope, hundreds of calls per job, under relentless and growing load. You will architect how we evaluate agentic Copilot experiences, from multi-turn trajectories and tool selection to end-to-end task outcomes, and set the standard for how the whole product defines and measures quality. This opportunity will allow you to shape the evaluation strategy for one of the world’s most visible AI products, work at the frontier of large-scale agentic systems and reliability engineering, and grow your influence as a technical leader across a large engineering organization. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. ResponsibilitiesSets the technical direction for the Copilot offline evaluation platform and partners with teams across the Copilot organization to turn ambiguous quality questions into rigorous, reproducible evaluations: the scenarios, metrics, and scorecards that gate what ships. Leads the architecture of an AI-first, agentic-first evaluation platform in which autonomous agents are first-class operators, designing the services, pipelines, and tooling that agents can run, extend, and reason about, with the observability and guardrails that make agent-driven operation trustworthy. Owns the platform’s hardest challenge: delivering trustworthy results across a deep chain of dependencies driven far outside their normal operating envelope, designing for graceful degradation, intelligent retry, dependency-aware gating, and reliability that holds under sustained, growing load. Defines how Copilot’s agentic experiences are evaluated end to end, from multi-turn trajectories and tool selection to task outcomes and modalities such as UX and voice, building the simulation, scraping, and scoring capabilities that make agent behavior measurable. Leads by example and mentors engineers across teams to produce extensible, maintainable systems, driving modernization of the evaluation...