Senior Site Reliability Engineer - India
Syfe · Gurgaon, Haryana, India · On-site
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Syfe is APAC's largest and fastest-growing digital wealth platform , trusted with over US$10 billion in assets. We are fundamentally changing how hundreds of thousands of people across Asia-Pacific build wealth through a holistic approach to managing money rather than just pushing investment products. Backed by world-class investors and recognised as a leader in wealthtech, we are a team of passionate builders creating the future of wealth management. About the Role :
We are looking for a Senior Site Reliability Engineer to own the reliability of Syfe's production platform end-to-end. Syfe
runs a Kubernetes-native, multi-region platform (Singapore, Hong Kong, Sydney) serving a regulated digital wealth-
management product, where availability, latency, and trust are first-order product features.
This is a senior individual-contributor role, not a managerial one. You will define what "reliable" means in measurable terms,
build the systems and automation that keep us there, and own and lead our on-call and incident-response program.
You'll spend your time engineering reliability into the platform — through SLOs, observability, and automation — rather than
firefighting, and you'll raise the bar for how the whole engineering org operates production
What You'll Own
Reliability targets. Define and drive SLIs/SLOs and error budgets across critical services; partner with product and
engineering teams to make error-budget-based decisions that balance velocity and stability.
On-call & incident program. Own the on-call rotation, escalation policies, and paging strategy. Establish incident
command, run blameless postmortems, and turn RCAs into tracked, completed reliability work. Drive down MTTD and
MTTR.
Platform & Kubernetes reliability. Own the reliability of our EKS-based deployment platform — GitOps delivery
(ArgoCD), Helm-based release configuration, and Infrastructure as Code (Terraform/OpenTofu) on AWS. Make
deployments safe, progressive, and reversible.
Observability. Build and mature the observability stack (Datadog for production APM/RUM; Grafana, VictoriaMetrics,
and ClickHouse for metrics and logs). Make systems debuggable: meaningful dashboards, actionable alerts, and low
alert noise.
Resilience. Lead capacity planning, scalability, failure-mode analysis, disaster-recovery and business-continuity
planning, and game-day / chaos exercises across regions.
Toil reduction. Identify operational toil and eliminate it with automation and self-service tooling, so reliability scales with
the platform rather than with headcount.
Production safety. Strengthen rollout/rollback paths, deployment guardrails, and secrets handling (HashiCorp Vault),
and partner with engineering teams to harden services before they reach production.
Engineering influence. Lead by example with hands-on engineering — design reviews, production-readiness reviews,
runbooks, documentation, and mentoring — embedding SRE practices across…