Senior AI Evaluation & Reliability Engineer
Aubergine · Ahmedabad / Remote · India · Remote
Pay: INR 3,000,000 – 3,500,000 a year
Posted Aug 27, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Why Aubergine
Aubergine is a global transformation and innovation partner , shaping next-gen digital products through consulting-informed execution that integrates strategy, design, and development.
Since 2013, we’ve built 400+ B2B and B2C products worldwide , turning powerful ideas into impact-driven experiences. We are one of the top global B2B companies on Clutch , rated highest among more than 80,000 technology service providers.
With more than 150 digital thinkers , we are home to some of the brightest, most passionate people around the world who are committed to delivering excellence.
We’re not just another workplace. We’re officially Great Place To Work® certified , with an exceptional trust index rating, making Aubergine a community where you can thrive and grow.
Role Overview: Build AI Systems We Can Trust
We are looking for a Senior AI Evaluation & Reliability Engineer who is passionate about solving one of the most important challenges in AI:
How do we know an AI system is actually working, improving, and delivering business value?
You will design and build production-grade evaluation systems for LLMs, RAG applications, and multi-agent systems, while helping enterprise clients and our engineering teams adopt AI with confidence.
This role goes beyond building evals. You will be a trusted AI consultant to clients, a technical mentor to engineers, and a key contributor to our journey towards becoming an AI superagency.
What You Will Own
Make AI Performance & ROI Measurable Build evaluation strategies that connect AI performance to business outcomes and ROI.
Define quality benchmarks, SLOs, risk thresholds, and success criteria.
Track hallucinations, reliability issues, and production risks.
Build executive-friendly AI quality and ROI scorecards.
Help clients determine where AI should be autonomous, supervised, or avoided.
Don't just measure model accuracy. Measure business impact.
Build Production-Grade Evaluation Pipelines Architect automated evaluation pipelines for LLMs, RAG, and agentic systems.
Integrate evaluations into CI/CD using GitHub Actions, GitLab CI, or equivalent.
Build regression suites for prompts, models, tools, and workflows.
Measure faithfulness, context precision, answer relevance, semantic drift, and task completion.
Establish statistically meaningful benchmarks and continuously monitor AI quality.
Evaluation should become part of engineering, not a final QA step.
Engineer LLM-as-a-Judge Systems Design reference-based and reference-free LLM evaluation frameworks.
Create structured rubrics and scoring systems.
Build calibration loops using human-labelled datasets.
Identify and mitigate judge biases such as position, verbosity, and self-preference bias.
Measure judge reliability and optimize evaluation quality, latency, and cost.
Own Evaluation Economics LLM evaluations can become expensive quickly.
You will:
Design tiered evaluation…