# Senior Research Engineer

> Weekday AI · Bengaluru, India · Full-time · Posted 2026-08-13

**Salary:** INR 5,000,000–9,000,000

**Workplace:** on_site

**Department:** Weekday's Client via platform

## Description

**This role is for one of Weekday’s clients  
Salary range: Rs 5000000 - Rs 9000000 (ie INR 50 - 90 LPA)**

Min Experience: 1+ years  
Location: Bengaluru  
JobType: full-time

We are looking for a highly motivated **Senior Research Engineer** with **1–8 years of experience** to join our research and engineering team. The ideal candidate will work at the intersection of **LLM research, model evaluation, and large-scale experimentation**, helping develop rigorous methods to understand, measure, and improve the capabilities of modern language models.

You will design and implement evaluation frameworks, build high-quality datasets and test suites, analyze model behavior, and translate research findings into practical improvements. This role is ideal for someone who enjoys solving open-ended research problems while also being comfortable building production-quality systems.

## Requirements

### Key Responsibilities

-   Design, develop, and maintain comprehensive **LLM evaluation (evals)** frameworks to measure model capabilities, reliability, reasoning, instruction following, safety, and task performance.
-   Conduct research on **large language models**, including model behavior, capabilities, limitations, prompting, fine-tuning, and evaluation methodologies.
-   Develop novel evaluation methodologies and experiments for emerging LLM capabilities and use cases.
-   Create high-quality evaluation datasets, test cases, rubrics, and automated evaluation pipelines.
-   Analyze model outputs using quantitative and qualitative methods to identify performance gaps and behavioral patterns.
-   Design controlled experiments to compare models, prompts, training approaches, and inference strategies.
-   Build scalable tooling for running evaluations across large numbers of prompts, models, and datasets.
-   Collaborate with researchers, ML engineers, and product teams to convert research insights into measurable model improvements.
-   Investigate failures and edge cases and develop targeted evaluations to capture previously undetected model weaknesses.
-   Contribute to technical documentation, research reports, internal benchmarks, and presentations of findings.
-   Stay current with developments in LLM research, evaluation techniques, reasoning systems, and AI benchmarks.

### Required Skills & Qualifications

-   **1–8 years of experience** in machine learning, AI research, software engineering, data science, or a related technical field.
-   Strong hands-on experience designing and implementing **LLM evals** or model evaluation systems.
-   Solid understanding of **LLM research**, including model capabilities, prompting, fine-tuning, inference, and evaluation methodologies.
-   Strong Python programming and experience working with ML/AI frameworks and data-processing pipelines.
-   Ability to formulate research questions, design experiments, interpret results, and communicate technical findings clearly.
-   Strong analytical and problem-solving skills with attention to experimental rigor and reproducibility.
-   Experience working with large datasets, automated testing, and evaluation pipelines.

### Good-to-Have Skills

-   Experience developing or working with **LLM benchmarks** and standardized evaluation suites.
-   Familiarity with benchmark design, dataset curation, scoring methodologies, and statistical analysis.
-   Experience with open-source LLMs, model APIs, Hugging Face, PyTorch, or similar frameworks.
-   Exposure to reinforcement learning, RLHF/RLAIF, fine-tuning, synthetic data generation, or agentic systems.
-   Research publications, technical blogs, open-source contributions, or demonstrated independent AI research work.

## Apply

[Apply at Weekday AI](https://apply.workable.com/weekday-1/j/0619619D6C/apply)

---
Powered by [Workable](https://www.workable.com)
