VP, Site Reliability Engineer
SGX · Singapore, SG · On-site
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
BUILD MARKETS. SHAPE ECONOMIES.
At SGX Group, we create markets. We turn ideas into products and expand access to new opportunities. We are building the future of the exchange in-house: from architecture and platforms to the critical systems that power markets. The biggest decisions are still open to people who join us now.
The opportunity
As Vice President, Site Reliability Engineering, you will lead reliability across our platforms and services and build SRE into a discipline the wider engineering organisation works to. You will own the standards for observability, automation, incident response, and resilience, along with the accountability for whether they hold when it matters.
Most engineering roles here manage the tension between agility and reliability. This one arbitrates it. Service levels, error budgets, and release decisions run through your remit, which means saying no sometimes and, harder, saying yes when the instinct is to hold. In market infrastructure, an outage is not an internal inconvenience. Everyone sees it.
There are strong foundations to build on, and a significant amount still to do. The engineering model, tooling, and ways of working will keep changing. If you want to inherit a mature SRE practice, this is not that one yet. You would be building it. It will suit someone who enjoys complexity, is comfortable leading through change, and wants to build capability that lasts.
The Site Reliability Engineering team
Site Reliability Engineering keeps the platforms behind SGX Group's business and market infrastructure available, observable, and recoverable.
The team sets the standards other engineering teams work to across service levels, error budgets, observability, incident response and automation. It works across engineering, infrastructure, security, and product, and it owns the practice as well as seen as the SME for ensuring resilience and proactive maintenance and continuous improvement of the estate.
The next phase is about establishing SRE as a discipline rather than a function, moving reliability decisions upstream into design, and reducing the manual work that currently sits behind keeping services up.
What you will do
Own the outcome
Define and lead the enterprise SRE strategy, establishing reliability engineering as a core discipline across critical products, platforms, and services.
Establish and govern SLO, SLI, and error budget practices, so reliability trade-offs are made on evidence rather than instinct.
Take accountability for availability, recovery readiness, and how services behave under load and under failure.
Set the engineering standard
Set the direction for observability, automation, self-healing, and resilience platforms, improving service stability, incident detection, and operational efficiency at scale.
Lead incident management, recovery readiness, post-incident learning, and chaos engineering, so resilience is tested rather than assumed.
Drive toil…