Sr. SRE Engineer - Dynatrace, Python, AIOPS

Bangalore, IndiaFull-timePosted Aug 6, 2026

About Chubb

Chubb is a world leader in insurance. With operations in 54 countries and territories, Chubb provides commercial and personal property and casualty insurance, personal accident and supplemental health insurance, reinsurance and life insurance to a diverse group of clients. The company is defined by its extensive product and service offerings, broad distribution capabilities, exceptional financial strength and local operations globally. Parent company Chubb Limited is listed on the New York Stock Exchange (NYSE: CB) and is a component of the S&P 500 index. Chubb employs approximately 40,000 people worldwide. Additional information can be found at: www.chubb.com.

About Chubb India

At Chubb India, we are on an exciting journey of digital transformation driven by a commitment to engineering excellence and analytics. We are proud to share that we have been officially certified as a Great Place to Work® for the third consecutive year, a reflection of the culture at Chubb where we believe in fostering an environment where everyone can thrive, innovate, and grow

With a team of over 2500 talented professionals, we encourage a start-up mindset that promotes collaboration, diverse perspectives, and a solution-driven attitude. We are dedicated to building expertise in engineering, analytics, and automation, empowering our teams to excel in a dynamic digital landscape.

We offer an environment where you will be part of an organization that is dedicated to solving real-world challenges in the insurance industry. Together, we will work to shape the future through innovation and continuous learning.

Position Details

  • Job Title: Sr. SRE Engineer - Dynatrace, Python,  AIOPS

  • Function/Department: Technology

  • Location: Hyderabad/Bangalore 

  • Employment Type: Full Time

  • Reports To:  TADEPALLI, MADHURI

 

Position Summary:

  • Act as a Senior, hands-on engineer within the SRE team — designing, building, and operating automation and self-healing systems that reduce toil and strengthen production reliability.
  • Take ownership of SLI/SLO instrumentation and error budget tracking for an assigned portfolio of services, partnering with development teams to drive measurable reliability improvements.
  • Serve as a senior on-call escalation point and incident responder, driving root cause analysis and durable hardening fixes for critical production issues.
  • Champion observability best practice through hands-on Dynatrace instrumentation, dashboard design, and alert tuning across owned services.
  • Mentor junior and mid-level SREs, raise engineering standards across the team, and contribute technical input to the team's automation and reliability roadmap.

Major Duties and Responsibilities

  • Technical Leadership & Mentorship
  • Act as a technical role model within the SRE team; set the standard for engineering quality, automation-first thinking, and reliability best practice.
  • Mentor and coach junior and mid-level SREs on troubleshooting techniques, automation approaches, and reliability engineering principles.
  • Review peers' designs, runbooks, and automation scripts; provide constructive feedback that raises the bar for engineering quality.
  • Contribute to technical interviews and skills assessment of SRE candidates when requested.
  • Support the SRE Lead in shaping the team's reliability roadmap by proposing, prototyping, and validating improvements.
  • Incident Response & On-Call
  • Participate in the 24/7 follow-the-sun on-call rotation as a senior escalation point (L2/L3); take ownership of complex, high-severity incidents.
  • Act as incident commander or technical lead for P1/P2 incidents when rostered; drive diagnosis, mitigation, and recovery under pressure.
  • Author and maintain blameless postmortems for incidents you own; track corrective actions through to closure.
  • Build and continuously improve runbooks and playbooks for services you support; automate manual steps wherever feasible — a runbook unchanged after 90 days is a toil backlog item.
  • Partner with Problem Management on root cause investigations; implement permanent fixes rather than workarounds.
  • Reliability Engineering & SLO Ownership
  • Define and instrument SLI/SLO measurements for assigned services; monitor error budget burn and flag risk to release velocity.
  • Build automation and self-healing tooling (Python, Go, or equivalent) that reduces manual toil and improves MTTD/MTTR for owned services.
  • Conduct capacity planning and performance analysis for owned services; recommend scaling and architecture improvements.
  • Participate in Production Readiness Reviews (PRR): assess new services against reliability criteria and provide sign-off recommendations.
  • Conduct Failure Mode and Effects Analysis (FMEA) for services you support; prioritise and implement hardening fixes.
  • Embed reliability practices into the SDLC for owned services — operability reviews, readiness checklists, and reliability testing as standard gates.
  • Chaos Engineering & Resilience
  • Design and execute chaos engineering experiments and GameDays for owned services; document findings and drive remediation.
  • Implement fault injection, load testing, and synthetic failure scenarios as part of CI/CD pipelines.
  • Apply resilience patterns — circuit breakers, graceful degradation, bulkhead isolation — to harden owned services.
  • Maintain and update the resilience scorecard for your service portfolio.
  • Observability & Dynatrace
  • Instrument applications and infrastructure using Dynatrace (OneAgent, OpenTelemetry) to achieve full-stack observability for owned services.
  • Build and maintain SLO dashboards, health scorecards, and Davis AI alerting rules for owned services.
  • Tune alert thresholds and reduce noise; continuously improve the signal-to-noise ratio for your service portfolio.
  • Contribute to team-wide observability standards: instrumentation guidelines, tagging taxonomy, and log retention policy.
  • AI-Augmented Operations
  • Build and maintain automation scripts and tooling that integrate with AIOps and LLM-based operational tools — incident summarisation, runbook generation, and knowledge retrieval.
  • Use AI-assisted triage and anomaly detection tools to accelerate diagnosis; provide feedback to improve model accuracy.
  • Pilot new AI/ML-based reliability tooling under guidance from the SRE Lead; document outcomes and recommendations.
  • Collaboration & Continuous Improvement
  • Partner with development and platform teams to influence service design for reliability and operability at design time, not post-deployment.
  • Contribute to quarterly SRE KPI reporting: SLO attainment, incident trends, and toil metrics for owned services.
  • Support knowledge transfer for project-to-support transitions; validate production readiness before go-live.
  • Stay current with SRE best practice (Google SRE principles, DORA metrics) and bring emerging techniques into the team.

 

  • Bachelor degree or higher in Computer Science, Software Engineering, Information Technology, or equivalent technical discipline.
  • Minimum 5 years of experience in software engineering, platform reliability, or SRE roles, with demonstrated hands-on ownership of production services.
  • Proven hands-on engineering capability: Python, Go, or equivalent; able to independently build automation tooling and self-healing workflows.
  • Practical experience implementing SLI/SLO frameworks and error budget policies for owned services.
  • Hands-on experience with observability platforms (Dynatrace preferred); able to instrument, dashboard, and tune alerting independently.
  • Experience participating in 24/7 on-call rotations and acting as a senior escalation point for critical incidents.
  • Experience conducting chaos engineering experiments or resilience/load testing.
  • Strong understanding of SDLC, Agile, and DevOps delivery models.
  • ITIL Foundation Certificate (desirable); SRE engineering skill and reliability outcomes take precedence over ITSM process credentials.

Leadership & Soft Skills

  • Strong individual contributor with a track record of mentoring junior and mid-level engineers.
  • Clear communicator: able to explain complex technical issues and trade-offs to both technical and non-technical stakeholders.
  • Calm and methodical under pressure; effective as a senior responder during critical production incidents.
  • Collaborative partner to development and product teams; builds trust and influence without formal authority.
  • Data-driven: uses SLO attainment, toil ratios, and MTTD/MTTR trends to prioritise own work and recommendations.
  • Self-directed and proactive: identifies reliability gaps and drives improvements without needing direction.
  • Growth mindset: continuously builds technical depth and stays current with SRE and reliability engineering practice.
  • Flexible and committed to on-call obligations and incident response, including weekends and public holidays when rostered.

Technical Skill

  • Observability & Monitoring
  • Dynatrace (proficient to expert): OneAgent instrumentation, Davis AI, SLO frameworks, synthetic monitoring, distributed tracing, RUM, log management, USQL.
  • Secondary platforms (working knowledge): Grafana, Prometheus, Splunk, ELK Stack, Azure Monitor, Log Analytics.
  • OpenTelemetry: instrumentation standards and telemetry pipeline basics.
  • Automation & Reliability Tooling
  • ServiceNow: working knowledge of Incident, Problem, Change, and CMDB modules sufficient to integrate SRE workflows.
  • Python, Go, or equivalent: production-grade automation tooling, self-healing workflows, and reliability engineering scripts.
  • CI/CD: Jenkins, Azure DevOps, GitHub Actions; implementing release gating and reliability checks in pipelines.
  • Application & Data
  • Enterprise application platforms: Java, .NET, or equivalent; working understanding of multi-tier and microservices systems.
  • Databases: MSSQL, MySQL, Oracle, Cosmos DB; SQL/PLSQL for diagnostics and trend analysis.
  • API and integration patterns: REST, SOAP, event streaming (Kafka, MQ), and service mesh observability.
  • Cloud & Infrastructure
  • Microsoft Azure (proficient preferred): App Services, AKS, Azure Monitor, Log Analytics.
  • Container and orchestration: Docker, Kubernetes; day-to-day operation and troubleshooting.
  • Infrastructure as Code: Terraform or Bicep for repeatable, auditable environment provisioning.
  • AI & AIOps
  • AIOps tooling: Dynatrace Davis AI, anomaly correlation, and automated root cause analysis.
  • LLM-based operational tooling: incident summarisation, runbook generation, and AI-assisted triage.
  • Scripting for intelligent automation: self-healing workflows and auto-remediation.

Desired

  • Insurance or financial services domain knowledge.
  • Dynatrace Associate or Professional certification.
  • Microsoft Azure Administrator or DevOps Engineer certification.
  • ITIL Foundation or higher certification.
  • Experience with chaos engineering frameworks (e.g. Chaos Monkey, Litmus, Gremlin).
  • Exposure to AI/ML pipelines or MLOps in a production reliability context.
  • Hands-on experience with AIOps tooling in an enterprise environment.
  • Familiarity with DORA metrics and industry reliability benchmarking.
  • Data interpretation and trend analysis skills across operational datasets.
  • Google Cloud Professional — DevOps Engineer or equivalent SRE certification.

 

Why Join Us?

  • Be at the forefront of digital transformation in the insurance industry.
  • Lead impactful initiatives that simplify claims processing and enhance customer satisfaction.
  • Work alongside experienced professionals in a collaborative, innovation-driven environment.

Why Chubb?

Join Chubb to be part of a leading global insurance company!

Our constant focus on employee experience along with a start-up-like culture empowers you to achieve impactful results.

  • Industry leader: Chubb is a world leader in the insurance industry, powered by underwriting and engineering excellence
  • A Great Place to work: Chubb India has been recognized as a Great Place to Work® for the years 2023-2024, 2024-2025 and 2025-2026
  • Laser focus on excellence: At Chubb we pride ourselves on our culture of greatness where excellence is a mindset and a way of being. We constantly seek new and innovative ways to excel at work and deliver outstanding results
  • Start-Up Culture: Embracing the spirit of a start-up, our focus on speed and agility enables us to respond swiftly to market requirements, while a culture of ownership empowers employees to drive results that matter
  • Growth and success: As we continue to grow, we are steadfast in our commitment to provide our employees with the best work experience, enabling them to advance their careers in a conducive environment

Employee Benefits

Our company offers a comprehensive benefits package designed to support our employees’ health, well-being, and professional growth. Employees enjoy flexible work options, generous paid time off, and robust health coverage, including treatment for dental and vision related requirements. We invest in the future of our employees through continuous learning opportunities and career advancement programs, while fostering a supportive and inclusive work environment. Our benefits include:

  • Savings and Investment plans: We provide specialized benefits like Corporate NPS (National Pension Scheme), Employee Stock Purchase Plan (ESPP), Long-Term Incentive Plan (LTIP), Retiral Benefits and Car Lease that help employees optimally plan their finances
  • Upskilling and career growth opportunities: With a focus on continuous learning, we offer customized programs that support upskilling like Education Reimbursement Programs, Certification programs and access to global learning programs.
  • Health and Welfare Benefits: We care about our employees’ well-being in and out of work and have benefits like Hybrid Work Environment, Employee Assistance Program (EAP), Yearly Free Health campaigns and comprehensive Insurance benefits.

Application Process

Our recruitment process is designed to be transparent, and inclusive.

  • Step 1: Submit your application via the Chubb Careers Portal.
  • Step 2: Engage with our recruitment team for an initial discussion.
  • Step 3: Participate in HackerRank assessments/technical/functional interviews and assessments (if applicable).
  • Step 4: Final interaction with Chubb leadership.

Join Us

With you Chubb is better. Whether you are solving challenges on a global stage or creating innovative solutions for local markets, your contributions will help shape the future. If you value integrity, innovation, and inclusion, and are ready to make a difference, we invite you to be part of Chubb India’s journey.

Apply Nowhttps://www.chubb.com/emea-careers/

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free