JobRaahGet matched free

Jobs

Principal Site Reliability Engineer

servicetitan · India Bengaluru, Karnataka · On-site

Posted Oct 1, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

Ready to be a Titan? We're looking for a Principal Site Reliability Engineer to join our Site Reliability & Infrastructure Engineering team. We run entirely on the cloud, and this team owns production intelligence, reliability, and operational automation — moving from reactive incident response toward systems that observe, correlate, reason, and act. We're building the next generation of SRE: not dashboards and runbooks, but agents and automation that make our platform self-aware and self-healing at scale. At Principal level, you won't just operate the platform — you'll define how we think about reliability across the entire engineering organization. You'll set the technical direction, establish the standards, and be the person other engineers turn to when the hardest infrastructure problems need to be solved. We make a huge impact on thousands of companies in the U.S. and abroad by enabling them to be more efficient and effective at running their business. We have a cultural foundation built on diversity, inclusion, and innovation, and we want you and your ideas to thrive at ServiceTitan. Come join us. What You'll Do Define the technical vision and long-term architecture for ServiceTitan's reliability and infrastructure platform — setting the standard for how we operate at scale. Lead the design and delivery of major platform initiatives: Kubernetes infrastructure, observability frameworks, SLO programs, and incident management systems. Partner with Engineering Managers, Staff Engineers, and product engineering leadership to review architecture and infrastructure decisions before they ship — and hold the bar on non-functional requirements across the org. Identify systemic reliability risks before they become incidents, and drive remediation across teams with org-wide impact. Define and operationalize SLIs, SLOs, and error budgets across ServiceTitan's platform — not just within the SRE team, but as a standard every engineering team adopts. Drive adoption of reliability and observability best practices across engineering — through documentation, design reviews, and direct partnership with product teams. Design and build AI-assisted operational systems — agents that correlate production telemetry, diagnose failure patterns, recommend or execute remediation, and participate safely in deployment decisions. Move routine investigation and triage from humans to automation, while preserving human judgment for high-risk changes. Partner with Infrastructure Engineering on deployment safety — progressive delivery, canary analysis, and AI-assisted promotion decisions that answer "is there enough production evidence to continue this deployment?" not just "did the deployment complete?" Mentor Staff and Senior SREs, raising the technical ceiling of the team through code reviews, architecture feedback, and direct pairing on complex problems. Contribute to the technical hiring bar — participate in system design interviews…