JobRaahGet matched free

Jobs

AI Test Engineer - VLabs

vialto · Bengaluru · India · On-site

Posted Jul 29, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About Vialto Labs (VLabs)   Vialto Labs (VLabs) is responsible for redesigning how work is delivered in the tax and immigration service lines, as well as driving operational efficiency across Vialto’s functional areas using AI. The team builds and deploys novel AI-enabled solutions that directly improve productivity and increase delivery quality for our clients. VLabs is accountable for rapidly turning innovative experiments into production-ready deliverables at scale and embedding them into day-to-day operations. This team focuses on the highest-impact workflows, creating standardized, repeatable capabilities that can be deployed globally. Operating with a mandate for speed and measurable outcomes, VLabs works alongside service line, product, and platform leaders.  About the Role   AI Test Engineering is a hands-on role within VLabs Quality Engineering, responsible for validating the performance, reliability, and integrity of AI-enabled solutions in production environments. This role operates at the intersection of AI engineering and quality assurance, ensuring that outputs from LLMs, OCR pipelines, document classification models, and agentic workflows perform as expected at scale and meet defined business performance thresholds.  Working closely with the Programme Test Manager and partnering with engineering, product, and delivery teams, this role translates AI testing strategy into executable frameworks, evaluation pipelines, and reusable assets embedded into the delivery lifecycle.  Success requires independent execution, strong technical depth, and the ability to proactively identify risks, patterns, and performance gaps while enabling rapid, production-grade deployment of AI capabilities.  Key Responsibilities   AI Evaluation & Test Design   Translate AI testing strategy into executable test scenarios across LLM outputs, document classification, extraction accuracy, agent workflows, and edge cases  Design adversarial and boundary test inputs to expose hallucination, misclassification, and failure modes  Validate AI outputs for structure, consistency, accuracy, and production readiness against defined performance thresholds  Evaluation Engineering & Automation   Build reusable Python-based evaluation frameworks, including output validation, hallucination detection, and scoring mechanisms  Develop parameterized test scripts reusable across features, models, and releases  Implement AI-as-Judge frameworks, including prompt design, scoring logic, and calibration of evaluation reliability  Embed evaluation frameworks into CI/CD pipelines to support continuous testing and deployment  Drift Detection & Quality Monitoring   Design and operate drift detection frameworks using fixed baseline datasets and scheduled re-evaluation  Establish thresholds to distinguish acceptable variation from performance degradation  Enable release gating by identifying regressions prior to production deployment  Ground Truth &…