Lead DevOps Engineer
StarHub Ltd · Petaling Jaya, MY · Malaysia · On-site
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Role Mission:
As a Lead DevOps Engineer at StarHub, you will own the end-to-end DevOps strategy, platform reliability, and release governance for mission-critical platforms. You will lead the design and evolution of CI/CD pipelines, Kubernetes platform architecture, cloud infrastructure, observability, and production operations practices that support high-volume telco customer journeys across provisioning, charging, billing, and network activation.
This role requires strong working knowledge of end-to-end telco flows, and hands-on exposure to CRM/BSS integrations, mediation, balance management, and rating or charging flows. You will work across engineering, operations, product, vendor, and business teams to ensure customer-impacting journeys are delivered safely, monitored effectively, and operated with a strong RCA and revenue-assurance mindset.
The role goes beyond execution. You will drive DevOps best practices, establish SRE capabilities, improve production stability, and guide engineers in delivering faster, safer, and more reliable platform changes across high-availability telco environments.
Responsibilities:
- Drive DevOps strategy and operating model for CRM/OM, and charging-adjacent platforms, acting as the technical escalation point for complex production issues.
- Ensure customer journeys across CRM, order orchestration, charging, billing, mediation, and network activation flows.
- Support and improve integrations with telco network and charging systems, including IN/OCS, PCRF/PCF, HLR/HSS/UDM, CRM/BSS, mediation platforms, and external partner APIs.
- Lead design reviews and operational readiness for flow design, order orchestration, service activation, service modification, suspension, restoration, termination, and charging-triggered workflows.
- Support production incidents, charging mismatches, balance or rating defects, revenue leakage investigations, reconciliation gaps, and high-severity operational escalations.
- Architect and govern end-to-end CI/CD pipelines using Jenkins, Pipeline-as-Code, GitLab CI, and GitOps practices such as Argo CD or Flux, enabling safe and repeatable releases.
- Lead trunk-based development adoption, automated testing, quality gates, deployment validation, rollback controls, and release safety mechanisms across microservices.
- Design and operate multi-cluster Kubernetes platforms such as EKS and OpenShift, including networking, RBAC, scaling, resilience, workload standardization, and platform governance.
- Govern AWS cloud infrastructure and Infrastructure as Code using Terraform and/or CloudFormation, ensuring security, scalability, cost efficiency, and operational resilience.
- Apply SRE practices by defining SLOs/SLIs, improving reliability and performance, leading incident management, and driving structured root-cause analysis and problem management.
- Establish observability and operational excellence using Splunk, Prometheus, Grafana, CloudWatch, ELK, alerting,…