This role is for one of the Weekday's clients
Salary range: Rs 500000 - Rs 1700000 (ie INR 5 - 17 LPA)
Min Experience: 6+ years
Location: Chennai, Bangalore, Hyderabad, Mumbai, Pune
JobType: full-time
We are looking for a highly skilled Senior Observability Engineer to design, implement, and scale enterprise-grade observability platforms that provide deep visibility into distributed systems. This role is ideal for professionals passionate about improving system reliability, performance, and operational excellence through modern observability practices.
As a key member of the engineering team, you will lead the design and deployment of comprehensive monitoring, logging, and tracing solutions while driving the migration from traditional monitoring platforms to a modern observability ecosystem. You will collaborate closely with platform, infrastructure, and application teams to establish best practices and improve the overall health and resilience of mission-critical systems.
Requirements
Key Responsibilities
- Design, develop, and manage scalable end-to-end observability solutions covering metrics, logs, traces, and alerting across enterprise environments.
- Lead the migration from legacy monitoring platforms to modern observability frameworks and cloud-native monitoring solutions.
- Deploy, administer, and optimize observability platforms running on Kubernetes or OpenShift environments.
- Build reusable dashboards, alerts, and monitoring standards to improve operational visibility and incident response.
- Develop and maintain Helm charts for deployment and lifecycle management of observability components.
- Implement automation for deployment, configuration management, and operational workflows using Python or Bash scripting.
- Collaborate with engineering and application teams to define observability standards and integrate monitoring into development workflows.
- Analyze system performance, identify bottlenecks, and recommend improvements that enhance platform reliability and scalability.
- Provide technical leadership, architectural guidance, and strategic recommendations for observability initiatives.
- Support production operations by troubleshooting complex monitoring and infrastructure issues.
- Contribute to continuous improvement initiatives and drive adoption of observability best practices across engineering teams.
Requirements
Must-Have Skills
- Strong hands-on experience with OpenTelemetry for instrumentation and telemetry collection.
- Expertise in the Grafana Enterprise Stack, including Mimir, Loki, and Tempo.
- Experience administering and scaling ITRS Geneos in enterprise environments.
- Strong knowledge of Prometheus and PromQL.
- Hands-on experience with Grafana, including dashboard creation, alerting, and data source management.
- Experience administering OpenShift or Kubernetes clusters.
- Expertise in developing and managing Helm Charts for Kubernetes deployments.
- Experience designing, deploying, and scaling enterprise observability platforms.
Good-to-Have Skills
- Experience with Google Cloud Observability or other cloud-native monitoring solutions.
- Automation and scripting experience using Python or Bash.
- Familiarity with CI/CD pipelines and enterprise deployment processes.
- Experience implementing observability in cloud-native or microservices-based architectures.
Preferred Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, or a related field.
- 8–15 years of experience in Platform Engineering, Site Reliability Engineering (SRE), DevOps, Infrastructure Engineering, or Observability Engineering.
- Proven experience implementing observability solutions at enterprise scale.
- Strong understanding of distributed systems, container platforms, and cloud-native technologies.
Soft Skills
- Strong analytical and problem-solving abilities.
- Strategic thinking with the ability to influence technical direction.
- Excellent communication and stakeholder management skills.
- Ability to collaborate effectively across cross-functional teams.
- Strong leadership, mentoring, and relationship-building capabilities.
- Service-oriented mindset with a focus on operational excellence.
- Ability to manage multiple initiatives in a fast-paced environment.