# (Seoul) Senior Site Reliability Engineer· Cancer Screening

> Lunit · Seoul, South Korea · Full-time · Posted 2026-09-14

**Workplace:** on_site

**Department:** Engineering

## Description

### "Conquering cancer through AI"

![](https://workablehr.s3.amazonaws.com/uploads/photos/465794/d00310aa650aceb8a49a6d850f65801c.jpg)

Lunit, a portmanteau of ‘Learning unit,’ is a medical AI software company devoted to providing AI-powered total cancer care.

Our AI solutions help discover cancer and predict cancer treatment outcomes, achieving timely and individually-tailored cancer treatment.

**🗨️ About the Team**

-   Lunit's Software Engineering (hereafter, "SE") department develops the Lunit INSIGHT product. Within the SE department, the Core Product Engineering team performs the Backend/Application development needed to productize INSIGHT AI models. We translate product requirements into development requirements and build Software as a Medical Device (SaMD) that complies with medical device guidelines.
-   Our team's core goal is to ensure the INSIGHT product has optimized performance and reliability so it can operate across diverse countries and deployment environments.

**🗨️ About the Position**

-   The Site Reliability Engineer role sits between development and operations, improving reliability, observability, automation, and operational systems so the INSIGHT product can run stably.
-   Rather than simply performing operational tasks, you will discover recurring problems and operational inefficiencies in production, analyze their root causes, and translate the findings into automation, improved observability, and improved operational processes.
-   You will collaborate with Lunit International's SRE team, understand differing operational environments and processes, and connect and align the global operating system with the domestic product development environment.

**🚩 Roles & Responsibilities**

You will design and improve the reliability, observability, deployment automation, and cloud infrastructure operations of the Lunit INSIGHT product to ensure it runs stably. This role is not limited to performing predefined operational tasks. You will identify gaps in the team's current SRE capabilities and operational systems, then define and drive the improvements needed.

-   Improve the stability, availability, performance, and operational quality of the Lunit INSIGHT product and related services.
-   Design and enhance cloud infrastructure, deployment, monitoring, and operational automation systems.
-   Analyze cloud resource usage and cost, and continuously drive cost optimization while considering stability and performance.
-   Diagnose, recover from, and perform root-cause analysis on incidents, and prevent recurrence through monitoring, alerting, runbooks, and automation.
-   Reliably manage and improve operational configurations and security elements such as per-environment/per-customer settings, certificates, secrets, and access permissions.
-   Understand the product domain and collaborate with development teams to build reliability in from the design and development stages.
-   Collaborate in English with Lunit International's SRE team, understanding and aligning differing operational environments and processes.
-   Build an on-call and incident-response system for 24/7 service operations, and participate in the actual on-call rotation to handle production incidents. As the team grows, evolve and lead a sustainable on-call operating model.

## Requirements

**🎯 Requirements**

-   5+ years of experience in SRE, DevOps, Platform Engineering, or production infrastructure operations
-   Experience directly designing and operating Azure-based production environments
-   Hands-on experience with Linux, networking, and containers
-   Experience designing or improving CI/CD, Infrastructure as Code, or operational automation systems
-   Experience analyzing production incidents based on monitoring, and performing recovery and recurrence-prevention activities
-   Experience collaborating with product and development teams to balance stability and development velocity
-   Ability to discuss and coordinate technical decisions fluently in English with overseas engineers and colleagues from diverse roles

**🏅 Preferred Experiences**

-   Experience establishing and improving the technical direction or operating systems of the SRE or Platform domain
-   Experience identifying the causes of recurring incidents or operational inefficiencies and leading structural improvements such as recurrence prevention or automation
-   Experience building or operating SRE practices such as on-call, incident management, postmortems, and SLO/SLA
-   Experience with Infrastructure as Code (Terraform, Bicep, ARM templates) and operational automation using Python, Bash, etc.
-   Experience operating container and deployment environments such as Kubernetes, GitHub Actions, and Azure DevOps
-   Experience using observability tools such as Azure Monitor, Application Insights, and Log Analytics, and incident-response/operations tools such as PagerDuty and ServiceNow

**📌 Tech Stack**

-   Scripting/Language: Python, Bash
-   Cloud/Infrastructure: Azure, Linux, Docker, Kubernetes, Terraform, Bicep
-   CI/CD: GitHub Actions, Azure DevOps Pipelines
-   Observability/Operations: Azure Monitor, Application Insights, Log Analytics, PagerDuty, ServiceNow, Custom Dashboards
-   Product Environment: Python, Go, FastAPI, PostgreSQL, Redis, RabbitMQ
-   Collaboration: Jira, Confluence, Slack, GitHub

**📝How to Apply**

-   CV/Resume in **English**, free format

-   Please note that an English resume is required for this position.
-   본 포지션은 영문 이력서 제출이 필수입니다.

**🏃‍♀️ Hiring Process**

-   Document Screening → Assignment → 1st Competency-based Interview (On-site) → 2nd Competency-based Interview (Online) → Culture-fit Interview (On-site) → Onboarding

-   Interviews will be conducted in English, and detailed instructions for the Assignment will be provided separately.
-   After the final interview, we may proceed with reference checks if needed.

**🤝 Work Conditions and Environment**

-   Work type: Full-time (3-month probation)
-   Work location : Lunit HQ(5F, 374, Gangnam-daero, Gangnam-gu, Seoul)
-   Salary: After negotiation  
    

🎸 ETC

-   If you misrepresent your experience or education or provide false or fraudulent information in or with your application, it may be grounds for cancellation of the employment.
-   Lunit is committed in providing the preferential processing to those eligible for employment protection (national merits and people with disabilities) relevant to related laws and regulations.

## Benefits

🌻 Benefits & Perks

-   The office is at a very convenient location, just a minute away from Gangnam Station Exit 3.
-   Meal Allowance is provided (up to 12,000 KRW per meal) when working at the office.
-   Latest computer models, such as Macs and 4K monitors are provided and can be renewed every three years.
-   Seminar registration fees and book purchases are covered.
-   Regular in-house AI and medical seminars are held.
-   In-house English lessons (aka Luniversal) are provided for English development.
-   Access to high-quality AI learning resources & deep learning DevOps system.
-   Up to 1.2 million KRW worth of benefits points can be claimed annually.
-   Holiday Allowances are provided in the form of gifts or vouchers for Korean National holidays, Seollal and Chuseok.
-   Congratulatory and Condolence allowances, along with paid time off are provided.
-   Annual medical checkups and employee accident insurance are provided.

## Apply

[Apply at Lunit](https://apply.workable.com/lunit/j/EFD208C780/apply)

---
Powered by [Workable](https://www.workable.com)
