Applied AI Engineer
ContactMonkey · Toronto · Canada · On-site
Pay: CAD 140,000 – 160,000 a year
Posted Oct 7, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Hey there! We're ContactMonkey 👋
Our mission? To power measurable employee engagement worldwide. And we'd love for you to join us!
About the job - Applied AI Engineer
Join the Engineering Team, where you'll help build our agentic future.
You'll work alongside senior engineers, Product, and our Chief Product & Technology Officer (CPTO) to design, prototype, and ship AI-powered capabilities quickly. This role is hands-on and iterative, focused on building production-grade agentic workflows that improve how internal communications are curated, designed, delivered, measured, and orchestrated.
This is an AI engineering role with real infrastructure ownership. You won't be handed a platform - you'll help build the one our AI features run on. If you like being close to both the model and the metal, this is that job.
Our stack, concretely: Ruby on Rails and Vue.js in a production SaaS codebase; Amazon Bedrock for model inference; AWS on EKS, provisioned with Terraform and Terragrunt across regions; Sidekiq for background work; MySQL and PostgreSQL. We're mid-migration to a GitOps deployment model with Argo CD and Karpenter. You'll touch most of this.
This is not "call an LLM and hope it works." It's careful system design, honest measurement, and shipping production systems that people rely on every working day.
This is a great fit if you've shipped LLM features in production, you're comfortable when the answer involves a Terraform plan rather than a prompt tweak, and you're ready to go deeper - with senior engineers around you to learn from.
Your impact
Agentic Design & Orchestration: Build agentic workflows and orchestration patterns (dynamic routing, tool-using agents, feedback loops) - contributing to the design and owning the implementation of well-scoped components.
Evals & Measurement: Own the eval harness for the features you ship. Build the datasets, the automated grading, and the offline regression suites that tell us whether a prompt or model change actually made things better. Raise the bar for evaluation across the team - judge design, offline and online metrics, and the judgment to know when a number is real. Turn production failures into evals that catch that class of failure next time.
Infrastructure for AI: Own work with SRE and build the infrastructure your features depend on, as code. Build the deployment path for AI workloads on EKS alongside our SRE and platform engineers.
Reliability, Cost & Safety: Keep our AI layer trustworthy as it grows. Instrument token spend, latency, and failure rates as first-class metrics. Design for the failure modes that matter - hallucination, timeout, rate limit, cost blowout - and make sure the system degrades gracefully instead of falling over. Respond when things behave unexpectedly in production.
Prompts & Model Configuration as Code: Treat prompts as versioned, reviewable, rollback-able artifacts rather than strings someone edited in a console. Own how we move a prompt…