AI Quality Engineer
Rootly · Toronto, Canada · Hybrid
Posted Jun 16, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
About Rootly
At Rootly , we are on a mission to be the go-to way companies respond when things go wrong, helping every organization be more reliable. We do this by building an industry-leading incident management platform that allows companies around the world to consistently and quickly resolve incidents. We are not simply transforming an industry, we are carving an entirely new +$B segment ourselves and need incredible talent to achieve this ambitious goal together.
Customers love Rootly. Some of the fastest growing companies around the world such as NVIDIA, Figma, Canva, Tripadvisor, Squarespace and more rely on Rootly to power their critical incident management process. They obsess over our delightful enterprise-ready platform and unique partnership model. See why our customers have reviewed us 5 stars on G2 .
Investors love Rootly. We are backed by some of the most respected funds in the world from Y Combinator to operators like the CTO of Dropbox and GitHub. We'd be happy to disclose our entire funding and profitability picture live during the interview. As a culture we relentlessly put transparency first. We conduct monthly financial reviews as a team so everyone has a pulse on the health of the business and publish what we are building in our weekly changelog .
Rootly is building the AI-native future of incident management, and we need someone who can push our AI to its limits before our customers do. As our AI Quality Engineer, you'll own the evaluation and optimization of Rootly's agentic AI features -- designing test scenarios, running adversarial prompts, interpreting outputs, and working directly with engineering and product to close the loop on performance.
This isn't traditional QA. You'll spend your days thinking like an attacker, a confused user, and a power user all at once -- probing how our AI agents reason, make decisions, and handle edge cases across complex incident workflows.
What you'll do
Design and execute prompt-based test scenarios that cover happy paths, edge cases, and adversarial inputs across Rootly's agentic AI features
Evaluate AI outputs for accuracy, relevance, consistency, and alignment with expected workflow behaviour
Build and maintain an evaluation framework; structured test libraries, scoring rubrics, and regression suites to track AI performance over time
Identify failure modes, hallucinations, reasoning gaps, and unexpected agent behaviours; document findings and work with engineers to resolve them
Partner with Product and Engineering on new AI feature releases, contributing to acceptance criteria and quality gates before launch
Define and track quality metrics (accuracy rates, failure frequency, regression trends) and report findings to stakeholders
Stay current on LLM evaluation techniques, prompt engineering best practices, and agentic testing methodologies
What we're looking for
+5 years in QA, product operations, AI/ML evaluation, or a closely related role
…