# Lead Site Reliability Engineer or Platform Engineer

> Weekday AI · Bengaluru, India · Full-time · Posted 2026-09-22

**Salary:** INR 5,000,000–10,000,000

**Workplace:** on_site

**Department:** Weekday's Client via platform

## Description

**This role is for one of Weekday’s clients  
Salary range: Rs 5000000 - Rs 10000000 (ie INR 50 - 100 LPA)**

  
Min Experience: 11+ years  
Location: Bengaluru, Karnataka, India  
JobType: full-time

We are seeking a highly experienced Lead DevOps/Platform (Native AI) Engineer to join the technology division of our organization. This role is responsible for the reliability, security, scalability, and engineering excellence of our cloud-native technology platform leveraging various platforms, systems and tooling.

You will lead infrastructure design, zero-touch deployment programs, observability and incident response, policy-as-code governance, and developer self-service — while advancing next-generation platform capabilities including governed automation, secure tool integrations for operational workflows, retrieval-augmented operational context, and production-grade evaluation and guardrails for emerging Agentic AI platform use cases.

The ideal candidate blends deep platform engineering knowledge with industry best practices, cutting-edge technology trends plus hands-on experience in one or few of latest AI capabilities. This is a technical leadership role: approximately 60% technical ownership, 30% stakeholder management, and 10% people leadership (when applicable).

## Requirements

**Responsibilities**

-   Design, build, and operate cloud-native infrastructure at scale: load balancers, containers, serverless, managed and self-hosted databases, caching, and supporting platform services.
-   Lead Infrastructure-as-Code programs — modular repositories, environment promotion, state strategy, and safe migration from monolithic to per-service or contract-based configurations.
-   Support CI/CD and release engineering including progressive delivery, rollback safety, and reliability gates tied to SLO and error-budget signals.
-   Automate operational workflows with code and event-driven orchestration platforms — emphasizing accuracy, repeatability, compliance, and security.
-   Partner on multi-cloud platform concerns: identity, edge/CDN, secrets management, and mobile telemetry pipelines.
-   Serve as escalation lead for P0/P1 incidents within assigned product POD scope; drive MTTI/MTTR improvement using severity- and complexity-aware measurement.
-   Advance intelligent incident response: contextual alert enrichment, retrieval over runbooks and operational context, schema-validated root-cause analysis, and approval-gated remediation.
-   Own on-call routing hygiene, on-call health, post-mortem closure, and permanent-fix tracking for assigned products.
-   Contribute to product reliability outcomes — feature breakage detection, synthetics, latency SLOs, and observability accuracy.
-   Implement enterprise observability standards: distributed tracing, structured logging, required service tags, and SLO-based alerting.
-   Build operational telemetry for automated and multi-step platform workflows — per-run correlation, failure taxonomy, and anomaly review.
-   Operate secure agent-tool connectivity (e.g., MCP-class protocols) for standardized integration with enterprise APIs and operational systems.
-   Support retrieval-augmented architectures over operational corpora with evaluation harnesses and retrieval drift detection.

Security, Compliance & Governance

-   Enforce security-by-default: just-in-time access, IAM least privilege, secrets management, and edge/WAF protections.
-   Implement policy-as-code at pull-request and deploy stages; drive configuration drift detection and closure SLAs.
-   Define permissions, guardrails, and autonomy zones for automated systems — with kill switches and immutable audit logs.
-   Stand up and operate an LLM/AI gateway for policy enforcement, safety scanning, cost attribution, and reliability-gated deployments where applicable.

Developer Experience & Platform Product Ownership

-   Treat the platform as a product: self-service catalog, environment stability, and toil reduction.
-   Lead assigned platform engineering initiatives end-to-end — milestones, KPI measurement, stakeholder demos, documentation, and knowledge transfer.
-   Serve as primary liaison to assigned Product PODs: sprint and release planning, SLA negotiation, status reporting, and platform syncs.
-   Develop runbooks, standards, and machine- and human-readable operational documentation.

Advanced Platform Capabilities

-   Design and operationalize multi-step automated workflows with explicit human-approval nodes, tool registries, and evaluation-in-CI gates.
-   Build platform APIs that wrap existing CI/CD and IaC operations — enabling safe automation without rip-and-replace.
-   Provision isolated execution environments for automated runs: no standing credentials, default-deny egress, per-run isolation, mandatory evaluation results.
-   Implement model routing and circuit breakers for cost efficiency and provider resilience where LLM-backed workflows are in scope.
-   Promote a governed platform delivery lifecycle with measurable exit criteria at each stage.

Leadership & Collaboration

-   Mentor engineers on incident command, IaC standards, observability hygiene, and safe adoption of platform automation.
-   Drive cross-functional alignment with Security and product development using innovative delivery practices and data-driven KPI reviews.
-   Contribute to hiring, onboarding, and knowledge-culture initiatives.
-   Validate system integrity, infrastructure designs, application deployments, and operational processes; recommend improvements.

Requirements

-   Bachelor's or Master's degree in Computer Science, Engineering, Software Engineering, or related field (or equivalent experience).
-   10+ years as a DevOps Engineer, Platform Engineer, SRE, or Infrastructure Engineer, with meaningful software or infrastructure development experience.
-   Production ownership of large, business-critical distributed systems (preferably container-based microservices).
-   Strong troubleshooting and analytical skills — ability to identify issues before they become client-impacting problems.
-   Strong communication skills — explain technical protocols, processes, and decisions to technical and non-technical stakeholders.
-   Prior technical lead or Pod Lead experience supporting multiple product teams.
-   Led platform migrations (IaC refactor, CI/CD modernization, observability uplift) with minimal production disruption.
-   Built production automated workflows with human-in-the-loop gates — not proof-of-concept only.
-   Demonstrated ability to stay current with industry trends, IT operations best practices, and emerging platform technologies.

Must-have skills

Terraform, AWS, agentic ai

Good-to-have skills

Kubernetes, docker, pyhton

## Apply

[Apply at Weekday AI](https://apply.workable.com/weekday-1/j/1572FAD9CF/apply)

---
Powered by [Workable](https://www.workable.com)
