Senior SRE

Ulaanbaatar, MongoliaFull-timePosted Aug 6, 2026

As a Senior Site Reliability Engineer (SRE), you will play a key role in designing, building, and maintaining highly available, secure, and scalable infrastructure across AND Global and its subsidiaries. You will collaborate closely with engineering teams to deliver reliable platform solutions, automate operational processes, and ensure system stability while supporting new business initiatives.

You'll join a collaborative team of four experienced engineers and contribute to shaping infrastructure best practices, operational excellence, and continuous improvement.

Key Responsibilities:

  • Monitor production systems, application availability, performance, and overall system - health.
  • Build and maintain reliable cloud and Kubernetes infrastructure.
  • Automate infrastructure provisioning, deployment, monitoring, and recovery processes.
  • Investigate production incidents, identify root causes, and implement permanent fixes.
  • Lead incident response and prepare post-incident reports.
  • Work with development teams to improve application reliability and release processes.
  • Define and maintain service-level indicators, service-level objectives, and availability targets.
  • Improve CI/CD pipelines and deployment procedures.
  • Perform capacity planning, performance tuning, and cost optimization.
  • Build dashboards, alerts, monitoring standards, and operational runbooks.
  • Maintain backup, disaster recovery, and high-availability procedures.
  • Identify technical risks, bottlenecks, and areas for continuous improvement.
  • Mentor junior engineers and support engineering best practices.

Requirements:

  • Bachelor’s degree in computer science, information technology, or equivalent practical experience.
  • At least 5 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or system administration.
  • Strong Linux administration and troubleshooting skills.
  • Production experience with AWS, Microsoft Azure, or Google Cloud.
  • Strong experience with Kubernetes and Docker.
  • Experience with infrastructure-as-code tools such as Terraform or OpenTofu.
  • Experience with CI/CD tools such as GitLab CI, GitHub Actions, Jenkins, or Argo CD.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Loki, ELK, OpenSearch, Datadog, or OpenTelemetry.
  • Ability to automate tasks using Python, Go, Bash, or another programming language.
  • Good understanding of networking concepts, including DNS, TCP/IP, HTTP/HTTPS, load balancers, VPNs, firewalls, and routing.
  • Experience supporting distributed applications, databases, storage, and messaging systems.
  • Experience with incident management, root-cause analysis, and postmortems.
  • Strong problem-solving, communication, and documentation skills.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free