JobRaahGet matched free

Jobs

Senior Site Reliability Engineer - India

Syfe · Gurgaon, Haryana, India · On-site

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

Syfe is APAC's largest and fastest-growing digital wealth platform , trusted with over US$10 billion in assets. We are fundamentally changing how hundreds of thousands of people across Asia-Pacific build wealth through a holistic approach to managing money rather than just pushing investment products. Backed by world-class investors and recognised as a leader in wealthtech, we are a team of passionate builders creating the future of wealth management. About the Role : We are looking for a Senior Site Reliability Engineer to own the reliability of Syfe's production platform end-to-end. Syfe runs a Kubernetes-native, multi-region platform (Singapore, Hong Kong, Sydney) serving a regulated digital wealth- management product, where availability, latency, and trust are first-order product features. This is a senior individual-contributor role, not a managerial one. You will define what "reliable" means in measurable terms, build the systems and automation that keep us there, and own and lead our on-call and incident-response program. You'll spend your time engineering reliability into the platform — through SLOs, observability, and automation — rather than firefighting, and you'll raise the bar for how the whole engineering org operates production What You'll Own Reliability targets. Define and drive SLIs/SLOs and error budgets across critical services; partner with product and engineering teams to make error-budget-based decisions that balance velocity and stability. On-call & incident program. Own the on-call rotation, escalation policies, and paging strategy. Establish incident command, run blameless postmortems, and turn RCAs into tracked, completed reliability work. Drive down MTTD and MTTR. Platform & Kubernetes reliability. Own the reliability of our EKS-based deployment platform — GitOps delivery (ArgoCD), Helm-based release configuration, and Infrastructure as Code (Terraform/OpenTofu) on AWS. Make deployments safe, progressive, and reversible. Observability. Build and mature the observability stack (Datadog for production APM/RUM; Grafana, VictoriaMetrics, and ClickHouse for metrics and logs). Make systems debuggable: meaningful dashboards, actionable alerts, and low alert noise. Resilience. Lead capacity planning, scalability, failure-mode analysis, disaster-recovery and business-continuity planning, and game-day / chaos exercises across regions. Toil reduction. Identify operational toil and eliminate it with automation and self-service tooling, so reliability scales with the platform rather than with headcount. Production safety. Strengthen rollout/rollback paths, deployment guardrails, and secrets handling (HashiCorp Vault), and partner with engineering teams to harden services before they reach production. Engineering influence. Lead by example with hands-on engineering — design reviews, production-readiness reviews, runbooks, documentation, and mentoring — embedding SRE practices across…