Site Reliability Engineer
๐ง๐ต๐ถ๐ ๐ฟ๐ผ๐น๐ฒ ๐ถ๐ ๐ณ๐ผ๐ฟ ๐ผ๐ป๐ฒ ๐ผ๐ณ ๐๐ต๐ฒ ๐ช๐ฒ๐ฒ๐ธ๐ฑ๐ฎ๐'๐ ๐ฐ๐น๐ถ๐ฒ๐ป๐๐
๐ฆ๐ฎ๐น๐ฎ๐ฟ๐ ๐ฟ๐ฎ๐ป๐ด๐ฒ: ๐ฅ๐ ๐ญ๐ฎ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ - ๐ฅ๐ ๐ฎ๐ด๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ (๐ถ๐ฒ ๐๐ก๐ฅ ๐ญ๐ฎ-๐ฎ๐ด ๐๐ฃ๐)
Experience: 6+ yrs
Location: Hyderabad, Bengaluru, Pune, Chennai, Tamil Nadu, India, Mumbai, Maharashtra, India
Job Type: Full-time
We are seeking an experiencedย Kafka Platform Engineerย with strong expertise in distributed systems, large-scale messaging platforms, and production operations. This role is ideal for professionals who are passionate about building, maintaining, and optimizing highly available streaming infrastructure while ensuring reliability, scalability, and operational excellence across enterprise environments.
As a Kafka Platform Engineer, you will be responsible for managing mission-critical messaging platforms, improving platform performance, automating operational processes, and supporting production environments. You will collaborate with infrastructure, application, DevOps, and engineering teams to deliver resilient streaming solutions, troubleshoot complex production issues, and continuously enhance platform reliability through automation, monitoring, and best practices.
Requirements
Key Responsibilities
- Design, deploy, manage, and optimize Apache Kafka clusters and large-scale messaging or streaming platforms.
- Monitor platform health, system performance, and resource utilization using modern monitoring, logging, and alerting tools.
- Troubleshoot production incidents, identify root causes, and implement long-term solutions to improve system stability.
- Automate operational tasks, deployments, and maintenance activities using scripting and infrastructure automation techniques.
- Collaborate with application and infrastructure teams to support messaging architecture, integrations, and production workloads.
- Implement performance tuning, capacity planning, and scalability improvements for distributed systems.
- Maintain high availability, fault tolerance, and disaster recovery strategies for messaging infrastructure.
- Ensure platform security, system compliance, and operational best practices across production environments.
- Develop operational documentation, runbooks, and knowledge-sharing resources to improve support efficiency.
- Participate in on-call support, incident response, and continuous improvement initiatives to maintain service reliability.
What Makes You a Great Fit
- 6+ years of experience managing distributed systems, production infrastructure, or platform engineering environments.
- Strong hands-on experience withย Apache Kafkaย or large-scale messaging and event streaming platforms.
- Deep understanding of distributed systems architecture, scalability, fault tolerance, and production operations.
- Experience with monitoring, logging, alerting, and observability tools for enterprise infrastructure.
- Proficiency in at least one scripting or programming language such asย Python,ย Bash, orย Java.
- Strong knowledge of Linux system administration, networking fundamentals, and troubleshooting methodologies.
- Experience automating operational workflows and improving platform reliability through scripting and infrastructure automation.
- Excellent analytical, debugging, and problem-solving skills with a proactive operational mindset.
- Strong communication and collaboration skills with the ability to work effectively across cross-functional engineering teams.
- Passion for building reliable, secure, and highly available platform infrastructure while continuously improving operational excellence.