# Senior Site Reliability Engineer

> Mozn · Cairo, Egypt (Remote) · Full-time · Posted 2026-08-17

**Workplace:** remote

**Department:** Focal

## Description

### **About Mozn**

MOZN is a leading Enterprise AI company enabling organizations to make informed decisions in two critical domains: Financial Crime Prevention and Enterprise Knowledge Intelligence.  
  
We’re a diverse, collaborative team of innovators united by a shared purpose: to build AI that delivers tangible business value, builds trust, and empowers people and organizations with augmented intelligence. Our culture is built on the relentless pursuit of excellence and meaningful impact.  
  
If you’re passionate about working alongside exceptional talent on world-class AI, and you want the autonomy and runway to do the best work of your career, join us in shaping the future of intelligent enterprises.

### **About the role**

We're hiring an AI Senior SRE: someone who carries a normal SRE workload — on-call rotation, incident response, root cause analysis, hands-on Kubernetes/cloud work — and builds the agentic layer that automates that workload over time. You are not exempt from operating production. You're the person best placed to know what should be automated, because you're the one doing it. This role exists because most reliability toil (triage, root-causing, remediation, SLO tracking, onboarding checks) is repetitive and well-defined enough to hand to an LLM-based agent with the right guardrails. Your job is to do the on-call/ops work like any SRE, then turn what you learn into agents — identify the workflow, build the agent, and earn the trust to let it act with increasing autonomy.

### **What you'll do**

-   Carry a normal on-call rotation and act as a hands-on responder: investigate, fix, and document incidents yourself, exactly like any SRE on the team — especially for workflows that don' t have an agent yet.
-   Go deep on application-level reliability, not just infrastructure: read and debug service code, understand business logic well enough to find the real root cause, and ship fixes or PRs directly into application repos when the fix belongs in the app, not the platform.
-   Design, build, and ship LLM-based agents (using tools like Claude Code, OpenAI Codex, or Kimi K2/K3) that plug into our existing stack: Kubernetes, cloud APIs, Prometheus/Grafana/ELK/Datadog, PagerDuty/Slack.
-   Define the "tool" interface each agent needs — the specific APIs, scripts, and read/write actions — and build safe wrappers around them.
-   Set clear guardrails for every agent: what it may do autonomously vs. what it must propose for human approval, with a bias toward human-in-the-loop until an agent has earned trust.
-   Own agent evaluation — define what "correct" and "safe" look like per agent, and build test/backtest suites against real historical incidents.
-   Continuously tune prompts, context, and tool schemas as an agent' s scope grows.
-   Partner with the SRE/platform team to find good agent candidates: repetitive, well-scoped, auditable workflows.
-   Report on agent impact — MTTD/MTTR/MTTX movement, false positive/negative rates, and engineer-hours of toil removed.
-   Keep a security- and compliance-first posture: audit trails for every autonomous action, least-privilege access to production, and alignment with Saudi data residency/regulatory requirements.

### Requirements

-   3+ years building production software with LLMs — agentic workflows, tool/function calling, multi-step planning, RAG — not just personal use of a chat assistant.
-   Hands-on experience shipping real work with an agentic coding tool such as Claude Code, OpenAI Codex, or Kimi K2/K3.
-   Strong Python (or similar) for building agent tooling, API wrappers, and orchestration.
-   Real, hands-on SRE experience: comfortable being a primary on-call responder, running incident response, and doing root cause analysis under pressure — not just familiar with the concepts.
-   Application-level debugging skill, not just infra: able to read a service' s codebase, trace a failure back to the actual line/logic causing it, and ship a fix yourself — SRE work here isn' t limited to restarting pods or scaling nodes.
-   Solid hands-on Kubernetes and cloud provider experience (AWS/GCP/OCI/Azure) and fluency with observability tools (Prometheus, Grafana, Datadog, ELK) — both as an operator and as integration points for agents.
-   Understands guardrails for autonomous systems: permissioning, approval gates, rollback paths, auditability.
-   Can build trust with technical stakeholders — every agent starts with limited autonomy and has to earn more.

### Nice to have

-   Experience in Saudi Arabia / MENA, ideally consulting across public or private sector clients.
-   Familiarity with Terraform/Ansible, Docker, VM/on-prem setups — useful context for the agents you' ll build, not the core job.
-   Experience building eval/backtest harnesses for LLM agents against historical incident data.
-   Background in ML engineering, LLMOps, or platform engineering.

### Benefits

-   You will be at the forefront of an exciting time for the Middle East, joining a high-growth rocket-ship in an exciting space
-   You will be given a lot of responsibility and trust. We believe that the best results come when the people responsible for a function are given the freedom to do what they think is best
-   The fundamentals will be taken care of: competitive compensation, top-tier health insurance, and an enabling culture so that you can focus on what you do best
-   You will enjoy a fun and dynamic workplace working alongside some of the greatest minds in AI
-   We believe strength lies in difference, embracing all for who they are and empowered to be the best version of themselves

## Apply

[Apply at Mozn](https://apply.workable.com/mozn-ai/j/9B9809BAAC/apply)

---
Powered by [Workable](https://www.workable.com)
