Senior Site Reliability Engineer (SRE)
Tubi - Canada · Toronto, Canada (Hybrid) · Hybrid
Pay: CAD 116,000 – 235,100 a year
Posted Sep 11, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
About the Role:
Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission is to engineer resilience from the ground up, enabling our product teams to innovate rapidly while ensuring our users have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through a culture of data-driven decision-making, blameless learning, and relentless automation.
As a Senior Site Reliability Engineer, you are a hands-on engineer who blends deep software development expertise with a passion for operational excellence. You will be responsible for designing, building, and running the resilient, scalable, and increasingly self-healing systems that power our products. You will apply sound engineering principles to solve our most complex reliability challenges, with a mandate to automate everything, eliminate toil, and write robust, maintainable code. You will be a force multiplier, mentoring other engineers and elevating the site reliability bar for the entire organization.
This is a hybrid role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week.
What You'll Do:
System Architecture & Design: Design, build, and maintain scalable, highly available, and fault-tolerant distributed systems. Partner with development teams as a reliability consultant, reviewing designs and influencing architectural decisions to ensure new services are built with reliability, observability, and performance as core principles, not afterthoughts.
Automation & Software Development: Write robust, performant, and maintainable code to automate operational tasks, and CI/CD pipelines. Build the internal tools, libraries, and frameworks that enable engineering teams to self-service their observability needs, reducing cognitive load and increasing their velocity.
Incident Response & Post-Mortem Analysis: Participate in a 24/7 on-call rotation, acting as a key technical leader and incident commander during critical service disruptions. Conduct deep, blameless root cause analyses (RCAs) that go beyond immediate fixes to identify and address systemic issues. Drive the implementation of corrective actions to prevent the recurrence of incidents.
Performance & Capacity Planning: Proactively monitor, measure, and optimize system performance to ensure low latency and high efficiency. Gather and analyze metrics from operating systems and applications to assist in performance tuning and fault finding. Analyze usage patterns and historical data to forecast capacity needs, ensuring our platform stays ahead of customer demand.
Building AI-Driven Automation: Building and integrating solutions that leverage our AIOps platform. This involves writing the code that consumes signals from…