JobRaahGet matched free

Jobs

Senior Site Reliability Engineer

Redwood Software · Canada · On-site

Pay: CAD 125,000 – 145,000 a year

Posted Sep 16, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric—with AI embedded across the automation lifecycle—Redwood accelerates the path to autonomous operations. Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT As a Senior Site Reliability Engineer (SRE) / DevOps Engineer you will be responsible for ensuring the stability, performance, scalability, and reliability of Redwood’s mission-critical SaaS platform. You will apply engineering principles to operational challenges, automate repetitive work, strengthen observability, and collaborate across teams to deliver a resilient and high-performing customer experience. Provide day-to-day management of system alerts, monitor system health, and escalate issues as necessary to maintain high availability. Participate in a 24x7, team-shared on-call rotation for critical SaaS platform incidents and provide support during emergencies. Lead incident response efforts to ensure fast and effective mitigation and resolution of production issues. Perform thorough Root Cause Analysis (RCA) and lead blameless post-mortems to identify systemic weaknesses and establish corrective actions that prevent recurrence. Collaborate with engineering teams to establish and enforce error budgets derived from Service Level Objectives (SLOs), balancing development velocity with system stability. Automate routine operational tasks to reduce manual effort and toil while increasing team efficiency. Design, deploy, and maintain cloud infrastructure using Infrastructure as Code (IaC), leveraging Terraform and Helm for deployment to EKS/Kubernetes clusters. Design, secure, and troubleshoot AWS cloud network architecture, including VPCs, subnetting, routing, security groups/NACLs, load balancers (ALB/NLB), and VPN/Transit Gateway connectivity across a multi-region, multi-account environment supporting EKS and Docker Swarm on EC2. Improve infrastructure health by developing and implementing checks, scripts, and automated remediation to proactively address known issues and enable platform self-healing. Maintain, develop, and evolve Continuous Integration/Continuous Delivery (CI/CD) deployment code and pipelines. Maintain existing infrastructure running on…