JobRaahGet matched free

Jobs

Quality Lead, Agentic AI Workflow Evaluation

Innodata Inc. · In Office - San Jose, California · United States · On-site

Pay: USD 75 – 85 a hour

Posted Sep 19, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers. Scope of the Role: We are standing up a dedicated onsite team to evaluate complex, real-world agentic AI workflows for a frontier AI customer. Reviewers work through ambiguous, multi-step scenarios inside isolated test environments, assessing whether AI agents complete tasks safely, respect user intent and consent, and hold up under close scrutiny. The Quality Lead is the person accountable for whether that output is any good. This is a senior individual contributor role. You will not manage the reviewers — that sits with the Engagement Manager — but you set the standard they are held to. You own the audit sample, run calibration, keep the rubric usable as real cases stress it, and train reviewers into the work. You are also the deputy: when the Engagement Manager is out, the engagement runs on you. The quality approach here is not fully defined. We expect you to build it in partnership with the customer's quality leads, or at minimum to take what they have, run it honestly, and come back with specific recommendations for where it falls short. What You’ll Own: Own the quality system for the engagement: audit design, sampling strategy, scoring standards, and how quality gets measured and reported Build that system with the customer's quality leads where none exists, and where one does, operate it and recommend concrete improvements based on what the data shows Re-score a sample of reviewer output as a second pass; identify error patterns rather than isolated mistakes Run calibration sessions: surface disagreement, work it to resolution, and document the reasoning so the outcome holds for future cases Maintain rubric health — flag criteria that are ambiguous, overlapping, or silent on cases the team keeps hitting, and drive revisions through the customer Train and onboard new reviewers, including nesting plans, ramp criteria, and the judgment call on when someone is production-ready Give the Engagement Manager the evidence behind performance conversations: who is drifting, on what, and whether coaching is working Report quality trends to the Engagement Manager and, alongside them, to the customer Deputize for the Engagement Manager on delivery operations during absences Maintain information security, privacy, and facility access practices required by the customer's onsite environment You’ll Thrive in This Role If You…