JobRaahGet matched free

Jobs

Site Reliability Engineer II (AI Platform)

OpenTable · Toronto, Canada · Hybrid

Pay: CAD 110,000 – 130,000 a year

Posted Sep 29, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

This hybrid role requires working in the office two days per week. With millions of diners, 70,000+ restaurant partners and 25+ years of experience, OpenTable, part of Booking Holdings, Inc. (NASDAQ: BKNG), is an industry leader with a passion for helping restaurants thrive. Our world-class technology empowers restaurants to focus on what matters most – their team, their guests, and their bottom line – while enabling diners to discover and book the perfect restaurant for every occasion. Every employee at OpenTable has a tangible impact on what we do and how we do it. You’ll also be part of a global team and its portfolio of metasearch brands. Hospitality is all about taking care of others, and it defines our culture. About the job As a Site Reliability Engineer II on the Serving Platforms team within Infrastructure Engineering, you will design, automate, and manage the core container stack and infrastructure powering our global business applications. Operating in a high-scale, self-hosted environment, you will serve as a subject matter expert for Kubernetes, Linux systems, and cloud-native automation, directly driving the reliability, security, and efficiency of our platform. In this role, you will collaborate with cross-functional engineering teams worldwide, lead greenfield infrastructure projects, resolve complex incidents, and build self-service capabilities that empower application developers across the organization. Responsibilities Maintain, tune, and ensure high availability for the low-level Linux operating system and Kubernetes control plane across our self-hosted bare-metal infrastructure. Architect, build, and maintain scalable container management, configuration management, and automation tools across global environments. Investigate, resolve, and conduct root-cause analysis for complex infrastructure disruptions and performance bottlenecks at the system call level. Participate in high-impact platform engineering projects and collaborate with globally distributed engineering teams to drive infrastructure standardization. Participate in the team's on-call rotation to support critical production systems and ensure operational resilience. Develop and maintain self-service tools, automation pipelines, and robust infrastructure monitoring to eliminate manual operational overhead. Minimum Qualifications 5+ years of hands-on Linux experience (e.g., Ubuntu, CentOS) with expertise in kernel tuning (sysctl), process management (cgroups/namespaces), system calls, and performance optimization. 3+ years of experience using configuration management systems such as Puppet, Chef, Ansible, or SaltStack in production environments. Proven experience building, operating, and troubleshooting bare-metal Kubernetes clusters from the ground up, including control plane, etcd, and CNI plugin management. Proficiency with continuous system automation and scripting in languages such as Go, Python, Ruby, Perl, or Bash. …