# Senior Operations Engineer

> Weekday AI · Bengaluru, India · Full-time · Posted 2026-08-04

**Salary:** INR 600,000–2,500,000

**Workplace:** on_site

**Department:** Weekday's Client via platform

## Description

**This role is for one of the Weekday's clients**

**Salary range: Rs 600000 - Rs 2500000 (ie INR 6 - 25 LPA)**

Min Experience: 4+ years

Location: Bengaluru  
JobType: full-time

We are seeking a highly skilled **Senior Operations Engineer** to support and optimize modern AI-powered production platforms. The ideal candidate will have strong expertise in **Python**, Site Reliability Engineering (SRE), DevOps, and GenAI application support. You will be responsible for ensuring the reliability, scalability, and operational excellence of AI workloads while collaborating with engineering teams to deploy and maintain production-grade Generative AI solutions.

This role is ideal for engineers who enjoy solving complex operational challenges, automating infrastructure, and supporting next-generation AI applications in cloud environments.

## Requirements

### **Key Responsibilities**

-   Manage, monitor, and optimize production environments supporting Generative AI applications and services.
-   Develop automation tools and operational utilities using **Python** to improve system reliability and operational efficiency.
-   Support deployment, maintenance, and troubleshooting of GenAI applications built using modern AI frameworks.
-   Monitor production systems, investigate incidents, perform root cause analysis, and implement preventive measures.
-   Collaborate with software engineering, platform, and infrastructure teams to ensure high availability and performance.
-   Deploy, maintain, and optimize AI workloads across **AWS**, **Google Cloud Platform (GCP)**, or similar cloud environments.
-   Support Retrieval-Augmented Generation (RAG) pipelines and AI application workflows.
-   Implement monitoring, logging, alerting, and incident response processes for production systems.
-   Optimize Linux-based infrastructure, networking, and system performance.
-   Contribute to CI/CD automation, infrastructure improvements, and operational best practices.
-   Maintain operational documentation, runbooks, and incident management procedures.

### **Requirements**

### **Must-Have Skills**

-   Strong programming expertise in **Python**.
-   **KARAT assessment completion** (mandatory, if applicable to the hiring process).
-   Hands-on experience developing or supporting **Generative AI (GenAI)** applications.
-   Experience building or supporting applications using **LangChain**, **LangGraph**, and **RAG (Retrieval-Augmented Generation)** pipelines.
-   Experience working with **AWS**, **Google Cloud Platform (GCP)**, or OpenAI platforms.
-   Strong experience in **Site Reliability Engineering (SRE)**, DevOps, Production Engineering, or Production Support.
-   Good understanding of Linux systems administration and networking concepts.
-   Experience with monitoring, troubleshooting, incident management, and production operations.
-   Strong understanding of cloud-native infrastructure and operational best practices.

### **Good-to-Have Skills**

-   Experience with **LangChain**.
-   Experience with **LangGraph**.
-   Knowledge of Kubernetes, Docker, or container orchestration platforms.
-   Familiarity with CI/CD pipelines and Infrastructure as Code (IaC).
-   Experience supporting Machine Learning or AI production workloads.
-   Exposure to observability tools, monitoring platforms, and log management solutions.
-   Knowledge of automation frameworks and scripting for infrastructure management.

### **Preferred Qualifications**

-   Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
-   4–10 years of experience in **SRE**, **DevOps**, **Production Engineering**, or **Platform Operations**.
-   At least 1 year of experience supporting **Generative AI** or **Machine Learning** production workloads is preferred.
-   Experience working in cloud-native, distributed, or AI-driven environments.

### **Soft Skills**

-   Strong analytical and troubleshooting skills.
-   Excellent problem-solving and critical thinking abilities.
-   Effective communication and cross-functional collaboration skills.
-   Ability to manage production incidents under pressure.
-   Strong ownership mindset with a focus on operational excellence.
-   Self-motivated, adaptable, and eager to learn emerging AI technologies.

### **Success Metrics**

Success in this role will be measured by maintaining high availability, reliability, and performance of production AI platforms while minimizing system downtime and incident resolution times. Performance will also be evaluated based on the successful deployment and operational stability of GenAI applications, the effectiveness of automation initiatives, and improvements in system monitoring and observability.

Additionally, success will be reflected in proactive incident prevention, continuous optimization of cloud infrastructure, efficient support for AI workloads, and strong collaboration with engineering teams to ensure secure, scalable, and reliable production environments.

## Apply

[Apply at Weekday AI](https://apply.workable.com/weekday-1/j/CA1300A323/apply)

---
Powered by [Workable](https://www.workable.com)
