Cloud DevOps -Operations Support-Puppet
Job Description
Site Reliability Engineer (SRE – SaaS Platform Operations)
Position Summary: Experienced Azure-based SRE required to support a large-scale enterprise SaaS platform, with strong Microsoft SQL Server expertise, future PostgreSQL readiness, production operations, patching, automation, troubleshooting, and incident response capability.
Role Type
Site Reliability Engineering / SaaS Platform Operations
Primary Technology Focus
Microsoft Azure, Microsoft SQL Server, PostgreSQL, Windows/Linux, Monitoring, Automation
Future Scope
Upcoming releases include database movement from MSSQL to PostgreSQL; PostgreSQL operational support is expected to become increasingly important.
Experience Level
4+ years relevant experience; strong MSSQL production support background; PostgreSQL exposure preferred
Working Model
24x7 production support environment; US time zone support/night shifts as required
Audience
Client, delivery leadership, hiring panel, and internal stakeholders
Position Overview
We are seeking an experienced Site Reliability Engineer (SRE) to support a large-scale enterprise SaaS platform operating in cloud and high-availability environments.
This role is responsible for maintaining infrastructure reliability, availability, performance, and operational excellence across Microsoft Azure, Microsoft SQL Server, PostgreSQL readiness, Windows, Linux, monitoring, automation, and incident response functions.
This is not a first-line helpdesk role. The role requires a hands-on engineer who can independently investigate complex technical issues, collaborate with engineering and platform teams, and provide clear technical communication to internal and client stakeholders.
Future Scope: In upcoming product releases, selected databases are expected to move from Microsoft SQL Server to PostgreSQL. The role should therefore include PostgreSQL awareness and operational readiness in addition to current MSSQL responsibilities.
Role Focus Areas
Focus Area
Expected Capability
Business Outcome
Cloud Operations
Manage and support Azure infrastructure, monitoring, storage, compute, and patching activities.
Stable, secure, and scalable platform operations.
Database Reliability
Administer MSSQL workloads today and support PostgreSQL readiness for future releases, including tuning, maintenance, backups, and HA/DR.
Improved database performance, resiliency, and recoverability.
Patching & Maintenance
Plan and support SQL/database patching and operating system patching across Windows and Linux environments.
Improved security compliance and reduced operational risk.
Incident Response
Perform initial analysis, support P1/P2 triage, contribute to RCA, and improve runbooks.
Reduced MTTR and stronger production readiness.
Automation & Observability
Use PowerShell/Python, alert tuning, logging, dashboards, and Infrastructure-as-Code practices.
Better operational efficiency and proactive issue detection.
Engineering Collaboration
Reproduce issues, validate defects, and escalate with evidence to product engineering.
Faster defect resolution and improved customer experience.
Key Responsibilities Infrastructure & Platform Reliability
Maintain highly available, reliable, and scalable cloud infrastructure in Microsoft Azure.
Monitor platform health, review technical logs, and proactively address performance and availability issues.
Improve infrastructure monitoring, alerting, and logging to support proactive reliability management.
Plan, coordinate, and support operating system patching and maintenance activities across Windows and Linux servers.
Ensure security updates, compliance patches, and platform upgrades are executed in line with change management processes while minimizing service disruption.
Drive operational excellence through standardization, automation, and continuous improvement.
Database Administration & Performance Management
Manage and support Microsoft SQL Server databases across production and non-production environments on Azure.
Support PostgreSQL operational readiness and future PostgreSQL database support as selected databases move from MSSQL to PostgreSQL in upcoming releases.
Troubleshoot and tune queries, SQL jobs, indexing, CPU, memory, I/O, and storage utilization.
Implement and maintain backup, restore, high availability, disaster recovery, and routine database maintenance processes.
Perform SQL Server and PostgreSQL database patching, upgrades, maintenance, and version lifecycle activities under enterprise change management processes.
Use T-SQL and scripts for investigations, reporting, and controlled production data fixes under change control.
Incident Management & Technical Troubleshooting
Investigate complex infrastructure, application, database, Windows, Linux, and Azure-related issues.
Participate in incident response, perform initial technical analysis, and contribute to root cause analysis documentation.
Reproduce customer-reported issues in staging or lab environments where required.
Escalate verified product defects to engineering teams with clear technical evidence, logs, and impact analysis.
Automation, DevOps & Observability
Create and maintain operational automation using PowerShell, Python, or equivalent scripting languages.
Support Infrastructure-as-Code and configuration management practices using tools such as Terraform and Puppet.
Collaborate on CI/CD and operational tooling improvements using platforms such as GitHub and Jenkins.
Enhance monitoring and alert management practices for SaaS product operations.
Stakeholder & Customer Communication
Collaborate closely with onsite DB SREs, application teams, platform teams, product support, and engineering.
Provide clear and timely communication on issue status, technical findings, risks, and next steps.
Join customer or onsite calls when required to explain technical findings and remediation actions.
Update runbooks, knowledge base articles, troubleshooting guides, and operational documentation.
Required Qualifications Education & Experience
Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Information Technology, or a related technical field.
4+ years of relevant experience in SRE, production operations, database administration, cloud infrastructure, or enterprise platform support.
2+ years of SaaS pipeline or SaaS product operations experience.
2+ years of hands-on experience managing Microsoft Azure cloud infrastructure and cloud monitoring tools.
4+ years in Microsoft SQL Server-centric roles such as DBA, Production Support Engineer, or Database SRE in 24x7 environments.
PostgreSQL exposure or willingness to support PostgreSQL environments as part of future platform modernization scope.
Experience supporting SQL/database patching and operating system patching activities in controlled production environments.
Willingness and ability to work night shifts or US time zone aligned support coverage when required.
Technical Skills Matrix
Skill Category
Required Skills
Priority
Database
MSSQL 2016+, T-SQL, query tuning, SQL jobs, stored procedures, indexing, backup/restore, HA/DR
Critical
PostgreSQL
PostgreSQL administration, performance tuning, backup/recovery, replication, HA awareness, and migration/coexistence readiness
Future Scope / High
Patching
SQL/database patching, OS patching, version lifecycle management, maintenance planning, and controlled change execution
High
Cloud
Microsoft Azure infrastructure, Azure VMs, Azure storage, Azure SQL or SQL on Azure VMs, Azure monitoring
Critical
Operating Systems
Windows Server administration and working knowledge of Linux servers
High
Monitoring & SRE
Application performance monitoring, alert management, logging, incident response, RCA, runbooks
High
Automation
PowerShell, Python, scripting for investigations, reporting, and operational automation
High
DevOps / IaC
Terraform, Puppet, GitHub, Jenkins, CI/CD awareness
Medium
Containerization
Docker, Kubernetes, and microservices architecture familiarity
Preferred
Preferred / Additional Qualifications
Experience setting up, configuring, and improving monitoring and tooling for SRE/DevOps operations of a SaaS product.
Experience supporting PostgreSQL databases in enterprise production environments.
Experience with database migration, modernization, or coexistence initiatives involving Microsoft SQL Server and PostgreSQL.
Knowledge of PostgreSQL performance tuning, backup/recovery, replication, and high availability architectures.
Experience with PagerDuty or similar incident notification and escalation platforms.
Experience with Elasticsearch is a strong advantage.
Ability to perform application debugging and collaborate effectively with Product Support, Platform, Database, and Engineering teams.
Ability to conduct research, evaluate technology options, make recommendations, and maintain a high level of expertise in systems software and operational tooling.
Strong understanding of compliance, policies, procedures, patching windows, and change control expectations in enterprise production environments.
Core Competencies
Independent technical troubleshooting
Customer-focused communication
Production incident ownership
Analytical thinking and RCA mindset
Cross-functional collaboration
Operational discipline and documentation
Hiring / Screening Guidance
Candidate screening should prioritize strong MSSQL production operations and Azure infrastructure experience, while also considering PostgreSQL exposure or readiness due to upcoming platform releases where selected databases are expected to move from MSSQL to PostgreSQL. The role is best suited for an engineer who can combine database reliability, cloud operations, patching, automation, and incident response in a large-scale SaaS environment.
Strong Fit Indicators
Potential Red Flags
MSSQL DBA / production support background
Pure DevOps profile with limited database depth
Azure infrastructure operations experience
NOC/helpdesk-only background
PostgreSQL exposure or database modernization experience
Database developer with no production operations
SQL and OS patching experience in controlled environments
Cloud-only profile without MSSQL troubleshooting
Hands-on incident/RCA ownership
No exposure to controlled patching/change management
PowerShell/Python automation
Limited ownership in incident response
Comfortable with US time zone support
No willingness to support future PostgreSQL scope
Role Value Proposition
This role provides the opportunity to contribute directly to reliability, performance, security, and operational maturity for a large-scale enterprise SaaS platform. The selected candidate will support the current Microsoft SQL Server estate, help prepare for future PostgreSQL adoption, and work with modern cloud, database, automation, patching, and monitoring technologies while collaborating closely with experienced SRE, database, platform, support, and engineering teams.
Job Description
Site Reliability Engineer (SRE – SaaS Platform Operations)
Position Summary: Experienced Azure-based SRE required to support a large-scale enterprise SaaS platform, with strong Microsoft SQL Server expertise, future PostgreSQL readiness, production operations, patching, automation, troubleshooting, and incident response capability.
Role Type
Site Reliability Engineering / SaaS Platform Operations
Primary Technology Focus
Microsoft Azure, Microsoft SQL Server, PostgreSQL, Windows/Linux, Monitoring, Automation
Future Scope
Upcoming releases include database movement from MSSQL to PostgreSQL; PostgreSQL operational support is expected to become increasingly important.
Experience Level
4+ years relevant experience; strong MSSQL production support background; PostgreSQL exposure preferred
Working Model
24x7 production support environment; US time zone support/night shifts as required
Audience
Client, delivery leadership, hiring panel, and internal stakeholders
Position Overview
We are seeking an experienced Site Reliability Engineer (SRE) to support a large-scale enterprise SaaS platform operating in cloud and high-availability environments.
This role is responsible for maintaining infrastructure reliability, availability, performance, and operational excellence across Microsoft Azure, Microsoft SQL Server, PostgreSQL readiness, Windows, Linux, monitoring, automation, and incident response functions.
This is not a first-line helpdesk role. The role requires a hands-on engineer who can independently investigate complex technical issues, collaborate with engineering and platform teams, and provide clear technical communication to internal and client stakeholders.
Future Scope: In upcoming product releases, selected databases are expected to move from Microsoft SQL Server to PostgreSQL. The role should therefore include PostgreSQL awareness and operational readiness in addition to current MSSQL responsibilities.
Role Focus Areas
Focus Area
Expected Capability
Business Outcome
Cloud Operations
Manage and support Azure infrastructure, monitoring, storage, compute, and patching activities.
Stable, secure, and scalable platform operations.
Database Reliability
Administer MSSQL workloads today and support PostgreSQL readiness for future releases, including tuning, maintenance, backups, and HA/DR.
Improved database performance, resiliency, and recoverability.
Patching & Maintenance
Plan and support SQL/database patching and operating system patching across Windows and Linux environments.
Improved security compliance and reduced operational risk.
Incident Response
Perform initial analysis, support P1/P2 triage, contribute to RCA, and improve runbooks.
Reduced MTTR and stronger production readiness.
Automation & Observability
Use PowerShell/Python, alert tuning, logging, dashboards, and Infrastructure-as-Code practices.
Better operational efficiency and proactive issue detection.
Engineering Collaboration
Reproduce issues, validate defects, and escalate with evidence to product engineering.
Faster defect resolution and improved customer experience.
Key Responsibilities
Infrastructure & Platform Reliability
Maintain highly available, reliable, and scalable cloud infrastructure in Microsoft Azure.
Monitor platform health, review technical logs, and proactively address performance and availability issues.
Improve infrastructure monitoring, alerting, and logging to support proactive reliability management.
Plan, coordinate, and support operating system patching and maintenance activities across Windows and Linux servers.
Ensure security updates, compliance patches, and platform upgrades are executed in line with change management processes while minimizing service disruption.
Drive operational excellence through standardization, automation, and continuous improvement.
Database Administration & Performance Management
Manage and support Microsoft SQL Server databases across production and non-production environments on Azure.
Support PostgreSQL operational readiness and future PostgreSQL database support as selected databases move from MSSQL to PostgreSQL in upcoming releases.
Troubleshoot and tune queries, SQL jobs, indexing, CPU, memory, I/O, and storage utilization.
Implement and maintain backup, restore, high availability, disaster recovery, and routine database maintenance processes.
Perform SQL Server and PostgreSQL database patching, upgrades, maintenance, and version lifecycle activities under enterprise change management processes.
Use T-SQL and scripts for investigations, reporting, and controlled production data fixes under change control.
Incident Management & Technical Troubleshooting
Investigate complex infrastructure, application, database, Windows, Linux, and Azure-related issues.
Participate in incident response, perform initial technical analysis, and contribute to root cause analysis documentation.
Reproduce customer-reported issues in staging or lab environments where required.
Escalate verified product defects to engineering teams with clear technical evidence, logs, and impact analysis.
Automation, DevOps & Observability
Create and maintain operational automation using PowerShell, Python, or equivalent scripting languages.
Support Infrastructure-as-Code and configuration management practices using tools such as Terraform and Puppet.
Collaborate on CI/CD and operational tooling improvements using platforms such as GitHub and Jenkins.
Enhance monitoring and alert management practices for SaaS product operations.
Stakeholder & Customer Communication
Collaborate closely with onsite DB SREs, application teams, platform teams, product support, and engineering.
Provide clear and timely communication on issue status, technical findings, risks, and next steps.
Join customer or onsite calls when required to explain technical findings and remediation actions.
Update runbooks, knowledge base articles, troubleshooting guides, and operational documentation.
Required Qualifications
Education & Experience
Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Information Technology, or a related technical field.
4+ years of relevant experience in SRE, production operations, database administration, cloud infrastructure, or enterprise platform support.
2+ years of SaaS pipeline or SaaS product operations experience.
2+ years of hands-on experience managing Microsoft Azure cloud infrastructure and cloud monitoring tools.
4+ years in Microsoft SQL Server-centric roles such as DBA, Production Support Engineer, or Database SRE in 24x7 environments.
PostgreSQL exposure or willingness to support PostgreSQL environments as part of future platform modernization scope.
Experience supporting SQL/database patching and operating system patching activities in controlled production environments.
Willingness and ability to work night shifts or US time zone aligned support coverage when required.
Technical Skills Matrix
Skill Category
Required Skills
Priority
Database
MSSQL 2016+, T-SQL, query tuning, SQL jobs, stored procedures, indexing, backup/restore, HA/DR
Critical
PostgreSQL
PostgreSQL administration, performance tuning, backup/recovery, replication, HA awareness, and migration/coexistence readiness
Future Scope / High
Patching
SQL/database patching, OS patching, version lifecycle management, maintenance planning, and controlled change execution
High
Cloud
Microsoft Azure infrastructure, Azure VMs, Azure storage, Azure SQL or SQL on Azure VMs, Azure monitoring
Critical
Operating Systems
Windows Server administration and working knowledge of Linux servers
High
Monitoring & SRE
Application performance monitoring, alert management, logging, incident response, RCA, runbooks
High
Automation
PowerShell, Python, scripting for investigations, reporting, and operational automation
High
DevOps / IaC
Terraform, Puppet, GitHub, Jenkins, CI/CD awareness
Medium
Containerization
Docker, Kubernetes, and microservices architecture familiarity
Preferred
Preferred / Additional Qualifications
Experience setting up, configuring, and improving monitoring and tooling for SRE/DevOps operations of a SaaS product.
Experience supporting PostgreSQL databases in enterprise production environments.
Experience with database migration, modernization, or coexistence initiatives involving Microsoft SQL Server and PostgreSQL.
Knowledge of PostgreSQL performance tuning, backup/recovery, replication, and high availability architectures.
Experience with PagerDuty or similar incident notification and escalation platforms.
Experience with Elasticsearch is a strong advantage.
Ability to perform application debugging and collaborate effectively with Product Support, Platform, Database, and Engineering teams.
Ability to conduct research, evaluate technology options, make recommendations, and maintain a high level of expertise in systems software and operational tooling.
Strong understanding of compliance, policies, procedures, patching windows, and change control expectations in enterprise production environments.
Core Competencies
Independent technical troubleshooting
Customer-focused communication
Production incident ownership
Analytical thinking and RCA mindset
Cross-functional collaboration
Operational discipline and documentation
Hiring / Screening Guidance
Candidate screening should prioritize strong MSSQL production operations and Azure infrastructure experience, while also considering PostgreSQL exposure or readiness due to upcoming platform releases where selected databases are expected to move from MSSQL to PostgreSQL. The role is best suited for an engineer who can combine database reliability, cloud operations, patching, automation, and incident response in a large-scale SaaS environment.
Strong Fit Indicators
Potential Red Flags
MSSQL DBA / production support background
Pure DevOps profile with limited database depth
Azure infrastructure operations experience
NOC/helpdesk-only background
PostgreSQL exposure or database modernization experience
Database developer with no production operations
SQL and OS patching experience in controlled environments
Cloud-only profile without MSSQL troubleshooting
Hands-on incident/RCA ownership
No exposure to controlled patching/change management
PowerShell/Python automation
Limited ownership in incident response
Comfortable with US time zone support
No willingness to support future PostgreSQL scope
Role Value Proposition
This role provides the opportunity to contribute directly to reliability, performance, security, and operational maturity for a large-scale enterprise SaaS platform. The selected candidate will support the current Microsoft SQL Server estate, help prepare for future PostgreSQL adoption, and work with modern cloud, database, automation, patching, and monitoring technologies while collaborating closely with experienced SRE, database, platform, support, and engineering teams.
An individual with min bachelor degree or PG. Well versed with client facing roles. Expect great communication and presentation skills. Hands on experience on multi cloud environment handling SRE / Devops / SQL DBA operations and support at an enterprise client environments.