JobRaahGet matched free

Jobs

Site Reliability Engineering Manager

Conifers.ai · Tel Aviv · Israel · On-site

Posted Sep 15, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

Conifers is transforming security operations centers (SOCs) with CognitiveSOC™ , its AI SOC platform, enabling enterprises and MSSPs to achieve SOC excellence. By leveraging agentic AI, Conifers helps security teams investigate complex, multi-tier incidents with speed, accuracy, and trust. Led by seasoned cybersecurity leaders and backed by SYN Ventures, PICUS Capital, and others, the company brings deep industry knowledge and innovation to an increasingly AI-driven threat landscape. We’re building an AI-native security platform that enables autonomous agents to investigate real-world threats at a massive scale. About The Role : At Conifers, reliability means more than keeping the infrastructure up. It means making sure our product works correctly, consistently, and predictably for our customers, 24/7. We are looking for a Site Reliability Engineering Manager to take end-to-end ownership of the health and reliability of our production system. This is a highly hands-on role. You will lead a small engineering team focused on production reliability and work closely with R&D, Product, and GTM to make sure the system is observable, measurable, stable, and ready for production. You will own the mechanisms that allow us to understand, at any point in time, whether the system is healthy and behaving as expected, from infrastructure and services to integrations, investigation pipelines, AI agents, and customer-facing functionality. When something goes wrong, you will help drive it from detection through investigation, resolution, and prevention. The goal is simple: make sure Conifers works reliably, accurately, and continuously for every customer. What You’ll Do : Own the end-to-end health and reliability of the Conifers production platform, beyond infrastructure availability. Build and continuously improve the monitoring, observability, alerting, dashboards, health checks, and operational tooling required to understand whether the system is working correctly. Define what "healthy" means across the platform and establish clear reliability, availability, performance, and product health metrics and targets. Proactively identify production issues, abnormal behavior, degradation, and reliability risks before they impact customers. Own and coordinate the response to production incidents and critical bugs, working with the relevant engineering teams until issues are fully resolved. Establish strong incident management, root cause analysis, postmortem, and follow-up processes to ensure we learn from failures and prevent recurrence. Work closely with R&D teams to improve system resilience, error handling, monitoring, scalability, and production readiness. Partner with Product to ensure new capabilities have clear production health indicators, monitoring, failure handling, and operational readiness before release. Work closely with GTM and customer-facing teams when production issues affect customers,…