Platform Engineer
LawZero · Montreal · Canada · On-site
Posted Aug 27, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Founded by Yoshua Bengio, LawZero is a nonprofit organization focused on AI safety. In charge of technology, the IT department oversees Cybersecurity, End User support, and the management of the compute environment used to achieve our mission.
You will own the platform layer that sits between our people and the GPU infrastructure: CI/CD pipelines, Kubernetes clusters, cloud environments, and the standards that keep them consistent and secure. This is a hands-on role on a small IT team, with a wide surface area and a lot of room to shape how things are built.
Key responsibilities
Design, deploy, and run Kubernetes clusters for research workloads, including autoscaling, network policies, and workload isolation, both on-premise and in the cloud.
Define, communicate, and enforce best practices for CI/CD pipelines: automated builds, tests, container image creation, and deployments, with security scanning and provenance built into the pipeline rather than bolted on.
Implement and manage internal services supporting our research and data teams, providing them with a reliable, secure platform. Examples include artifacts registry, Container registries, etc.
Manage our cloud environments as code (Terraform or equivalent): accounts, networking, identity, secrets, and cost visibility.
Define and champion platform standards. Base images, deployment patterns, environment promotion, so researchers and engineers ship without reinventing the plumbing each time.
Work at the boundary between Kubernetes and our HPC cluster: containerized workflows that need to interface with the ones running on Slurm, shared storage access, and tooling that makes both environments feel coherent to a researcher.
Build observability into the platform: metrics, logs, traces, and alerts that make failures obvious and debugging quick. We currently use Prometheus and Grafana.
Embed security controls into everything above: least-privilege IAM, secrets management, supply-chain integrity for dependencies and images, network segmentation, and audit trails.
Document what you build and automate what you repeat. Reduce the number of things that only work because one person remembers how.
Participate in incident response for platform services, and in the post-incident work that keeps the same problem from recurring.
Skills and qualifications
Experience
3–5 years in platform engineering, DevOps, SRE, or a closely related infrastructure role.
Production experience with Kubernetes. Not just deploying to it, but operating it: upgrades, RBAC, networking, storage, troubleshooting a cluster that is misbehaving.
Solid CI/CD experience with a modern toolchain (GitHub Actions, GitLab CI, Jenkins, or similar), including building pipelines from scratch.
Experience architecting, deploying, and maintaining a GitOps workflow is an asset.
Hands-on experience with at least one major cloud provider (AWS, GCP, or Azure) and infrastructure as code. Familiarity with…