# Inference Engineer

> Weekday AI · Bengaluru, India · Full-time · Posted 2026-07-29

**Salary:** INR 2,400,000–3,600,000

**Workplace:** on_site

**Department:** Weekday's Client via platform

## Description

𝗧𝗵𝗶𝘀 𝗿𝗼𝗹𝗲 𝗶𝘀 𝗳𝗼𝗿 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗪𝗲𝗲𝗸𝗱𝗮𝘆'𝘀 𝗰𝗹𝗶𝗲𝗻𝘁𝘀

𝗦𝗮𝗹𝗮𝗿𝘆 𝗿𝗮𝗻𝗴𝗲: 𝗥𝘀 𝟮𝟰𝟬𝟬𝟬𝟬𝟬 - 𝗥𝘀 𝟯𝟲𝟬𝟬𝟬𝟬𝟬 (𝗶𝗲 𝗜𝗡𝗥 𝟮𝟰-𝟯𝟲 𝗟𝗣𝗔)

Experience: 5+ yrs

Location: Bengaluru, Karnataka, India

Job Type: Full-time

We are seeking a highly skilled **Inference Engineer** with strong expertise in **Large Language Models (LLMs)** and **vLLM** to build, optimize, and scale high-performance AI inference platforms. This role is ideal for professionals who are passionate about deploying production-grade generative AI systems, optimizing model serving, and improving inference efficiency for large-scale applications.

As an Inference Engineer, you will be responsible for designing, implementing, and maintaining scalable LLM inference infrastructure that delivers low-latency, high-throughput AI experiences. You will work closely with AI Researchers, Machine Learning Engineers, Platform Engineers, and Product teams to deploy state-of-the-art language models, optimize serving pipelines, and improve resource utilization across GPU-based environments. This role offers the opportunity to work with cutting-edge AI technologies while solving complex performance, scalability, and infrastructure challenges.

## Requirements

### Key Responsibilities

-   Design, deploy, and maintain scalable inference infrastructure for Large Language Models (LLMs) in production environments.
-   Build and optimize high-performance model serving pipelines using **vLLM** and other modern inference frameworks.
-   Improve inference latency, throughput, memory utilization, and overall system performance across GPU clusters.
-   Deploy, monitor, and manage open-source and proprietary LLMs while ensuring reliability and scalability.
-   Optimize GPU resource allocation, batching strategies, caching mechanisms, and model parallelism for efficient inference.
-   Collaborate with Machine Learning Engineers to transition trained models into production-ready inference services.
-   Develop APIs, microservices, and deployment workflows for AI-powered applications.
-   Implement monitoring, logging, benchmarking, and performance profiling for inference workloads.
-   Troubleshoot production issues related to model serving, infrastructure, and distributed inference systems.
-   Contribute to the continuous improvement of AI platform architecture, automation, and deployment best practices.

### What Makes You a Great Fit

-   5+ years of experience in Machine Learning Infrastructure, AI Platform Engineering, MLOps, or Inference Engineering.
-   Strong hands-on expertise with **Large Language Models (LLMs)** and **vLLM** for production-scale inference.
-   Experience deploying and optimizing transformer-based models using modern inference frameworks.
-   Solid understanding of GPU computing, CUDA, distributed inference, model quantization, and performance optimization techniques.
-   Experience with Python and deep learning frameworks such as PyTorch, Hugging Face Transformers, or similar ecosystems.
-   Knowledge of containerization and orchestration technologies including Docker and Kubernetes.
-   Familiarity with cloud platforms such as AWS, Azure, or GCP for AI infrastructure deployment.
-   Strong understanding of REST APIs, microservices, distributed systems, and scalable backend architectures.
-   Experience implementing monitoring, benchmarking, and observability for AI inference workloads.
-   Excellent analytical, debugging, and problem-solving skills with the ability to optimize complex AI systems.
-   Strong collaboration and communication skills, with the ability to work effectively across AI, platform, and product engineering teams.
-   Passion for advancing generative AI infrastructure and delivering reliable, high-performance inference solutions for real-world applications.

## Apply

[Apply at Weekday AI](https://apply.workable.com/weekday-1/j/A44098974C/apply)

---
Powered by [Workable](https://www.workable.com)
