Single Position
View All JobsSite Reliability Engineer 2
India, Telangana, HyderabadApply nowFind out how well you match with this jobJob number200042743Date postedJul 28, 2026Work site3 days / week in-officeTravelLess than 25%ProfessionSoftware EngineeringDisciplineSite Reliability EngineeringRole typeIndividual ContributorEmployment typeFull-TimeOverviewMicrosoft is a company where passionate innovators come to collaborate, envision what can be and take their careers further. This is a world of more possibilities, more innovation, more openness, and the sky is the limit thinking in a cloud-enabled world.
Microsoft’s Azure Data engineering team is leading the transformation of analytics in the world of data with products like databases, data integration, big data analytics, messaging & real-time analytics, and business intelligence. The products our portfolio include Microsoft Fabric, Azure SQL DB, Azure Cosmos DB, Azure PostgreSQL, Azure Data Factory, Azure Synapse Analytics, Azure Service Bus, Azure Event Grid, and Power BI. Our mission is to build the data platform for the age of AI, powering a new class of data-first applications and driving a data culture.
Within Microsoft Fabric, the Data Integration team enables organizations to connect, move, and shape data across their entire data estate. As data is generated across applications, devices, and systems, our integration capabilities make it easy to ingest, transform, and prepare data for downstream use. Fabric’s unified integration experience simplifies how customers bring data together, ensuring it is trusted, usable, and ready to power analytics, real-time insights, and AI scenarios.
The Customer Data Integration (CDI) organization within Power Query is building a new Live site engineering team that serves as the first line of defense for one of Microsoft's most critical data integration services. You will own incident triage and response across Power BI, Fabric, Power Query Online, Gateway, and hundreds of data connectors, and you will build the agentic automation that makes that operation increasingly self-driving.
This is not a passive monitoring role. You will be on call, triaging real incidents, and ensuring uptime for business-critical services. You will simultaneously build the intelligent agents, automation pipelines, and tooling that reduce manual effort with every iteration. Your goal is to make yourself more effective over time by engineering your way out of repetitive work.
We do not just value differences or different perspectives. We seek them out and invite them in so we can tap into the collective power of everyone in the company. As a result, our customers are better served.
Responsibilities
- Incident triage and first-line response: Provide on-call coverage for incoming incidents across CDI services. Perform initial investigation, severity assessment, and routing to owning engineering teams.
- Agentic triage system development: Build and extend AI-driven agents that ingest ICM alerts, correlate with recent deployments and feature flag rollouts, check known-issue databases, and produce initial assessments with suggested severity and owning team.
- TSG and known-issue matching: Develop automation that matches incoming incidents to relevant Troubleshooting Guides (TSGs) and known issues across Fabric and Power Platform — reducing investigation time and enabling faster resolution.
Auto-routing and classification: Configure and extend ICM routing rules and build intelligent classification systems based on service tree, alert signatures, and historical patterns.
- Incident lifecycle automation: Build agents for incident summarization, customer communications drafting, postmortem generation, and reporting, replacing manual authoring with AI-assisted workflows requiring human judgment only for high-severity incidents.
- Metrics and continuous improvement: Measure and improve first-time mitigation (FTM) rates, incident deflection rates, and time-to-triage. Use data to identify patterns, propose new TSGs, and drive systemic improvements.
- CSS quality bar: Defend the standard for incoming escalations by enforcing organizational and process standards
- Embody our culture and values
Qualifications
Required/Minimum Qualifications
- 4+ years of software engineering experience in site reliability, Live site operations, or incident management for cloud services.
- Good programming skills in one or more of: C#, PowerShell, Python, KQL/Kusto.
- Experience with incident management systems and workflows (ICM, PagerDuty, ServiceNow, or similar).
- Experience with monitoring, alerting, and observability systems (Kusto, Geneva, Grafana, or similar).
- Ability to work in an on-call rotation across time zones in a geographically distributed team.
- Experience interface with engineers, leadership, support, and customers.
Job Requirements: Other & Additional
- Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings: Microsoft Cloud Background Check:
- This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.
Preferred/Additional Qualifications
- Experience building AI/ML-driven automation, agents, or intelligent workflows (e.g., using LLMs, Copilot extensibility, MCP servers, or agentic frameworks).
- Familiarity with Live site ecosystem management (including log traversal, incident management, telemetry analysis, etc.)
- Experience with Azure, Power BI, and Fabric services.
- Experience with Troubleshooting Guide (TSG) authoring and incident pattern analysis.
- Understanding of SLA management, customer communications, and escalation workflows for cloud services.
Benefits/perks listed below may vary depending on the nature of your employment with Microsoft and the country where you work.
#azdat
#azuredata
#dataintegration #powerbi #powerquery #dataflows #fabric
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
Insights from previous hires