# Senior Site Reliability Engineer

> Kody · Hong Kong, Hong Kong · — · Posted 2026-09-08

**Workplace:** on_site

**Department:** Technology

## Description

### **Job Summary**

Kody is seeking a **Senior Site Reliability Engineer (8+ years of experience)** to drive the reliability, availability, scalability, and operational excellence of our global payment platform. Based in **Hong Kong or Shenzhen**, you will take end-to-end ownership of production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating across Europe, Asia, and North America.

### **Key Responsibilities**

-   **Incident Management & On-Call:** Participate in a follow-the-sun production on-call rotation as a senior incident responder. Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectiveness.
-   **Production Operations:** Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
-   **SLO & Reliability Engineering:** Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed services.
-   **Continuous Optimization:** Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA).
-   **Security & Compliance:** Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environments.
-   **Technical Leadership:** Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices.

## Requirements

### **Qualifications & Requirements**

-   **Experience:** **8+ years** of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systems.
-   **Core Technical Stack:** Strong expertise in **AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux**, networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana).
-   **Distributed Systems Mastery:** Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration.
-   **Domain Expertise:** Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standards.
-   **SRE Methodology:** Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reduction.
-   **Location & Communication:** Based in **Hong Kong or Shenzhen**. Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams.

### **Leadership & Operational Excellence**

-   **Ownership:** Demonstrates strong end-to-end accountability for service reliability and customer impact under high pressure.
-   **Structured Problem Solving:** Applies a systematic and data-driven approach to troubleshooting, telemetry analysis, and incident resolution in complex distributed environments.
-   **Crisis Management:** Proven ability to command cross-functional incident response efforts, align stakeholders, and maintain clear communication during critical outages.
-   **Engineering Culture:** Champions a blameless post-incident culture, operational readiness, continuous learning, and technical mentorship.

## Benefits

\- Competitive Package

\- A dynamic and innovative team

\- Collaborative, inclusive working environment

## Apply

[Apply at Kody](https://apply.workable.com/kody/j/17017C53C3/apply)

---
Powered by [Workable](https://www.workable.com)
