Inference Engineer

Bengaluru, IndiaFull-timePosted Jul 29, 2026

๐—ง๐—ต๐—ถ๐˜€ ๐—ฟ๐—ผ๐—น๐—ฒ ๐—ถ๐˜€ ๐—ณ๐—ผ๐—ฟ ๐—ผ๐—ป๐—ฒ ๐—ผ๐—ณ ๐˜๐—ต๐—ฒ ๐—ช๐—ฒ๐—ฒ๐—ธ๐—ฑ๐—ฎ๐˜†'๐˜€ ๐—ฐ๐—น๐—ถ๐—ฒ๐—ป๐˜๐˜€

๐—ฆ๐—ฎ๐—น๐—ฎ๐—ฟ๐˜† ๐—ฟ๐—ฎ๐—ป๐—ด๐—ฒ: ๐—ฅ๐˜€ ๐Ÿฎ๐Ÿฐ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ - ๐—ฅ๐˜€ ๐Ÿฏ๐Ÿฒ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ (๐—ถ๐—ฒ ๐—œ๐—ก๐—ฅ ๐Ÿฎ๐Ÿฐ-๐Ÿฏ๐Ÿฒ ๐—Ÿ๐—ฃ๐—”)

Experience: 5+ yrs

Location: Bengaluru, Karnataka, India

Job Type: Full-time

We are seeking a highly skilledย Inference Engineerย with strong expertise inย Large Language Models (LLMs)ย andย vLLMย to build, optimize, and scale high-performance AI inference platforms. This role is ideal for professionals who are passionate about deploying production-grade generative AI systems, optimizing model serving, and improving inference efficiency for large-scale applications.

As an Inference Engineer, you will be responsible for designing, implementing, and maintaining scalable LLM inference infrastructure that delivers low-latency, high-throughput AI experiences. You will work closely with AI Researchers, Machine Learning Engineers, Platform Engineers, and Product teams to deploy state-of-the-art language models, optimize serving pipelines, and improve resource utilization across GPU-based environments. This role offers the opportunity to work with cutting-edge AI technologies while solving complex performance, scalability, and infrastructure challenges.

Requirements

Key Responsibilities

  • Design, deploy, and maintain scalable inference infrastructure for Large Language Models (LLMs) in production environments.
  • Build and optimize high-performance model serving pipelines usingย vLLMย and other modern inference frameworks.
  • Improve inference latency, throughput, memory utilization, and overall system performance across GPU clusters.
  • Deploy, monitor, and manage open-source and proprietary LLMs while ensuring reliability and scalability.
  • Optimize GPU resource allocation, batching strategies, caching mechanisms, and model parallelism for efficient inference.
  • Collaborate with Machine Learning Engineers to transition trained models into production-ready inference services.
  • Develop APIs, microservices, and deployment workflows for AI-powered applications.
  • Implement monitoring, logging, benchmarking, and performance profiling for inference workloads.
  • Troubleshoot production issues related to model serving, infrastructure, and distributed inference systems.
  • Contribute to the continuous improvement of AI platform architecture, automation, and deployment best practices.

What Makes You a Great Fit

  • 5+ years of experience in Machine Learning Infrastructure, AI Platform Engineering, MLOps, or Inference Engineering.
  • Strong hands-on expertise withย Large Language Models (LLMs)ย andย vLLMย for production-scale inference.
  • Experience deploying and optimizing transformer-based models using modern inference frameworks.
  • Solid understanding of GPU computing, CUDA, distributed inference, model quantization, and performance optimization techniques.
  • Experience with Python and deep learning frameworks such as PyTorch, Hugging Face Transformers, or similar ecosystems.
  • Knowledge of containerization and orchestration technologies including Docker and Kubernetes.
  • Familiarity with cloud platforms such as AWS, Azure, or GCP for AI infrastructure deployment.
  • Strong understanding of REST APIs, microservices, distributed systems, and scalable backend architectures.
  • Experience implementing monitoring, benchmarking, and observability for AI inference workloads.
  • Excellent analytical, debugging, and problem-solving skills with the ability to optimize complex AI systems.
  • Strong collaboration and communication skills, with the ability to work effectively across AI, platform, and product engineering teams.
  • Passion for advancing generative AI infrastructure and delivering reliable, high-performance inference solutions for real-world applications.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your rรฉsumรฉ so you only see the matches that matter.

Get started free