# Senior Site Reliability Engineer

> Weekday AI · Bengaluru, India · Full-time · Posted 2026-09-19

**Salary:** INR 1,300,000–2,000,000

**Workplace:** on_site

**Department:** Weekday's Client via platform

## Description

𝗧𝗵𝗶𝘀 𝗿𝗼𝗹𝗲 𝗶𝘀 𝗳𝗼𝗿 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗪𝗲𝗲𝗸𝗱𝗮𝘆'𝘀 𝗰𝗹𝗶𝗲𝗻𝘁𝘀

𝗦𝗮𝗹𝗮𝗿𝘆 𝗿𝗮𝗻𝗴𝗲: 𝗥𝘀 𝟭𝟯𝟬𝟬𝟬𝟬𝟬 - 𝗥𝘀 𝟮𝟬𝟬𝟬𝟬𝟬𝟬 (𝗶𝗲 𝗜𝗡𝗥 𝟭𝟯-𝟮𝟬 𝗟𝗣𝗔)

Experience: 4+ yrs

Location: Bengaluru, Karnataka, India

Job Type: Full-time

We are looking for an experienced **Senior Site Reliability Engineer (SRE)** to build, operate, and continuously improve highly reliable, scalable, secure, and high-performing production systems across **hybrid and multi-cloud environments**.

The role combines cloud infrastructure, Kubernetes, automation, observability, incident management, and reliability engineering. The ideal candidate will have strong hands-on experience with **AWS, Kubernetes, Terraform, Python, Bash, and modern observability platforms**, along with a strong understanding of production operations and distributed systems.

## Requirements

Key Responsibilities

-   Define and manage **SLIs, SLOs, SLAs, error budgets, and reliability objectives** for critical production services.
-   Drive initiatives to improve system availability, scalability, performance, resilience, and operational efficiency.
-   Manage and support production **Kubernetes environments**, including Amazon EKS and Red Hat OpenShift.
-   Deploy and maintain containerised workloads using **Docker, Kubernetes, and Helm**.
-   Manage cloud infrastructure across **AWS and IBM Cloud**, including hybrid-cloud environments.
-   Design and maintain reliable cloud connectivity, networking, disaster-recovery, and failover solutions.
-   Develop and maintain infrastructure using **Terraform and Infrastructure as Code (IaC)** practices.
-   Automate operational processes, infrastructure tasks, and troubleshooting workflows using **Python and Bash**.
-   Build and enhance observability solutions using **Prometheus, Grafana, OpenTelemetry, Thanos**, and logging platforms.
-   Monitor system health, identify performance bottlenecks, and proactively address reliability risks.
-   Participate in and lead **high-severity incident response** and production troubleshooting.
-   Conduct root-cause analysis and lead post-incident reviews and corrective actions.
-   Develop and maintain capacity-planning and reliability-improvement strategies.
-   Implement secure, resilient, and compliant infrastructure practices across cloud environments.
-   Support disaster-recovery planning, testing, and continuous improvement.
-   Collaborate with software engineering, platform, security, and architecture teams to improve production reliability.
-   Contribute to architecture reviews, engineering standards, operational best practices, and automation initiatives.
-   Mentor engineers and promote strong SRE, DevOps, observability, and production-engineering practices.

What Makes You a Great Fit

-   **4–6 years of professional experience** in Site Reliability Engineering, DevOps, Cloud Infrastructure, or a closely related field.
-   Strong hands-on experience with **AWS and Kubernetes** in production environments.
-   Experience managing **Amazon EKS, Docker, and Helm**.
-   Practical experience with **Red Hat OpenShift** is highly desirable.
-   Strong proficiency in **Terraform** and Infrastructure as Code practices.
-   Hands-on scripting and automation experience using **Python and Bash**.
-   Strong experience with **Prometheus and Grafana** for monitoring and observability.
-   Experience with **OpenTelemetry, Thanos, logging platforms**, or similar observability technologies.
-   Strong understanding of **SLIs, SLOs, error budgets, incident management, and production troubleshooting**.
-   Good understanding of DNS, TCP/IP networking, TLS, VPNs, load balancing, firewalls, and cloud connectivity.
-   Experience working with hybrid or multi-cloud infrastructure, preferably including **AWS and IBM Cloud**.
-   Strong understanding of containers, distributed systems, scalability, availability, and fault tolerance.
-   Experience with disaster recovery, capacity planning, and production resilience.
-   Exposure to regulated or compliance-driven environments such as **HIPAA, SOC 2, PCI DSS, or ISO 27001**.
-   Strong analytical, troubleshooting, and root-cause analysis skills.
-   Excellent communication and collaboration skills.
-   Ability to take ownership of critical production systems and operate effectively during high-severity incidents.
-   Experience mentoring engineers and contributing to technical architecture and reliability standards.

## Apply

[Apply at Weekday AI](https://apply.workable.com/weekday-1/j/2210AECEF2/apply)

---
Powered by [Workable](https://www.workable.com)
