Senior Principal Site Reliability Engineer
Bybit · Kuala Lumpur, Malaysia · On-site
Posted Aug 7, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
About Us
Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 — connecting users to the future of digital finance.
Our core values define how we build. We listen, care and improve to create products and experiences that put users first. Backed by a global team of ambitious builders, problem-solvers, and innovators, we foster a high-performance and fast-moving environment where talent is empowered to drive real impact at the global scale. Supported by 24/7 multilingual customer service and a strong commitment to innovation, we are shaping the future of finance through technology, collaboration, and bold execution.
Today, Bybit is recognized as one of the most trusted and transparent platforms in the digital asset industry, continuing to expand its global presence while building the infrastructure for the next generation of financial services.
Core Responsibilities
Chaos Engineering Platform Architecture & Development (50%)
Design and build an enterprise-grade chaos engineering platform supporting multi-cluster (K8s + EC2 hybrid), multi-region, and multi-environment (testnet/mainnet) deployments
Core capability development:
Fault Injection Engine: Pod-level / Node-level / AZ-level fault simulation, network latency / packet loss / partition, dependency timeout / error injection
Production Safety Assurance: Blast radius control, one-click Kill Switch, automatic rollback, real-time impact monitoring
Traffic Isolation: Experiment traffic tagging and isolation to ensure fault injection does not impact real users
Fault Isolation: Precise impact scoping at service / cluster / AZ granularity
Design experiment orchestration capabilities supporting complex fault scenario composition (e.g., simultaneous network latency + downstream timeout + cache invalidation)
Deep integration with existing monitoring, alerting, and SLO systems to achieve an automated closed loop: inject fault → observe impact → determine pass/fail
Production Resilience Validation Framework (30%)
Define safety standards and approval workflows for mainnet fault injection
Design and drive routine chaos experiments:
Daily patrol-level experiments: Low-risk experiments executed automatically on a daily/weekly basis
Periodic validation experiments: Monthly/quarterly resilience verification of critical paths
Large-scale drills: Cross-AZ / cross-region disaster recovery failover validation
Establish a resilience scoring system to quantify system health based on experiment results
Deliver improvement recommendations and drive business teams to remediate identified weaknesses
3. Technology Selection & Team Enablement (20%)
Evaluate and…