Site Reliability Engineering Lead, Product Enablement
otppb · Toronto, Canada · On-site
Posted Oct 5, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
The opportunity
The Site Reliability Engineering Lead (SREL) is accountable for the reliability, health, and operational sustainability of a broader portfolio of applications, data pipelines, platforms, and services. The role proactively identifies systemic and recurring risks and leads or directs technical initiatives, including the design and implementation of monitoring, automation, operational support tooling, and selected code to reduce incidents, manual intervention, cost, and operational burden. The SREL serves as a senior technical advisor and escalation point, translating complex issues, options, risks, and trade-offs into clear business language for stakeholders.
Who you'll work with
Reports to: Director, Product / Data Enablement
Works closely with: Site Reliability Engineering Specialists, Product Engineering Leads (PEL), Data Solutions Leads, Data Platform Leads, Enterprise Architecture, Security, Technology Services, business stakeholders, and third-party vendors.
Collaborates regularly with Product Engineering, Data Solutions, and Data Platform teams to support smooth transitions from build to run and to identify and resolve recurring reliability and supportability issues. Depending on the nature and complexity of the change, the Lead may implement or direct selected code, configuration, automation, or process changes through development, testing, release, and post-implementation validation, or work with the appropriate team to secure prioritization, ownership, and completion.
Partners with Technology Services teams including Infrastructure, Cloud, and Platform teams to troubleshoot complex and cross-service issues and to design and implement, or direct the implementation of, observability, automation, resilience, and operational support tooling that improves service quality and reduces operational burden.
Frequent communication with business stakeholders is required to provide transparency into system health, incident status, systemic risks, and improvement priorities, adapting the level of technical detail to the audience, including senior business stakeholders when required.
What you'll do
Service Reliability Strategy and Ownership
Own and continuously improve the reliability, supportability, and operational efficiency of a broad portfolio of applications, data pipelines, platforms, and services including availability, latency, performance, resilience, and cross-service dependencies.
Define and govern Service Level Availability (SLAs), and related reliability metrics for the portfolio in accordance with enterprise standards and business expectations and drive changes where performance or operational risk warrants.
Establish and maintain a prioritized reliability roadmap incorporating application, data pipeline and service health plans balancing business criticality, operational risk, support effort, capacity, cost, and technology priorities
Use deep technical and trend analysis across incidents,…