# Senior Software Engineer (Performance)

> Qode · Vietnam, Vietnam · Full-time · Posted 2026-07-23

**Workplace:** on_site

## Description

Job Description:

We are looking for a Senior Inference Engineer with a strong foundation in software engineering, distributed systems, and performance optimization to build and optimize inference engines for large-scale LLM serving systems. You will work across both research and production environments, ensuring our LLM serving systems are fast, scalable, and efficient. The role spans the entire inference stack — from kernel and runtime to scheduling, memory management, and distributed execution

  

Key Responsibilities:

-   Profile, benchmark, and analyze bottlenecks for LLM inference workloads across multiple layers: kernel, memory, networking, and scheduler
-   Optimize inference engines (vLLM, SGLang, TensorRT-LLM) for throughput, latency, memory efficiency, GPU utilization, and cost
-   Implement and fine-tune inference optimization techniques including batching, KV-cache management, quantization, speculative decoding, parallelism strategies, and disaggregated serving
-   Build instrumentation and profiling tools to identify bottlenecks
-   Ensure the reliability of the inference pipeline through A/B launches, rollback, model versioning, and fault tolerance
-   Collaborate with the Platform Engineering team to improve serving architecture based on performance findings
-   Document and share knowledge, contributing to internal best practices and AI open-source projects whenever possible

Requirements

1 - Mandatory:

-   At least 5 years of experience as a Software Engineer, Performance Engineer, or equivalent.
-   Strong foundation in Software Engineering, Software Architecture, and Distributed Systems.
-   Proficiency in at least one of the following languages: Python, Go, or C++.
-   Experience developing or optimizing distributed systems, high-throughput backends, or large-scale serving systems.
-   Experience with benchmarking, profiling, and performance tuning in production environments.
-   Ability to analyze CPU, Memory, Network, or Storage bottlenecks.
-   Strong systems thinking, Root Cause Analysis capabilities, and the ability to solve complex performance problems.
-   Strong ownership mindset and the ability to work independently.

2 - Nice to Have:

-   Experience with Linux internals, kernel tuning, or custom Linux kernel.
-   Understanding of GPU Architecture or CUDA Programming.
-   Experience with AI/ML Serving Systems or LLM Inference.- Have worked with one of the inference engines such as vLLM, SGLang, TensorRT-LLM, or Triton Inference Server.
-   Understanding of batching, KV Cache, quantization, speculative decoding, tensor/pipeline parallelism, or disaggregated serving.
-   Experience with the NVIDIA inference stack (TensorRT, Triton, CUTLASS, NCCL, cuBLAS, cuDNN).
-   Experience with observability stacks such as Prometheus, Grafana, or OpenTelemetry.
-   Open-source contributions or research related to AI Infrastructure, ML Systems, or Performance Optimization.

## Apply

[Apply at Qode](https://apply.workable.com/qodeworld/j/62BE2EAA11/apply)

---
Powered by [Workable](https://www.workable.com)
