Lead Site Reliability Engineer/ Expert
SITA Switzerland Sarl · Cluj-Napoca, UNAVAILABLE, RO · Romania · On-site
Posted Oct 5, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Description External
WELCOME TO SITA
At SITA, we keep airports moving, airlines flying smoothly, and borders open. Our technology and communication innovations power the success of the global air travel industry.
You'll find us in 95% of international airports, working closely with over 2,500 transportation and government clients. Each partnership brings unique challenges, and we thrive on delivering fresh solutions and cutting-edge tech to keep operations running like clockwork. We don't just move the world forward-we're proud to be recognized as a Great Place to Work ® by 79% of our employees and certified in most of our growing locations. Here, we feel empowered, supported, and inspired to grow.
Are you ready to love your job?
The adventure begins right here, with you, at SITA.
ABOUT THE ROLE AND THE TEAM
As Lead Site Reliability Engineer/ Expert you will be responsible for the proactive support of products so that there is high product performance that is continuously improved. Responsible for identifying and resolving the root causes of operational incidents, implementing solutions to improve stability and prevent recurrence. Manages the creation and maintenance of the event catalog to trigger events and develop both manual remediation approaches and automated workflows to resolve alerts. Oversee the deployment of IT services and solutions ensuring successful integration with minimal disruption. Focuses on operational automation and integration to enhance efficiency and collaboration between development and operations within service operations.
WHAT YOU WILL DO:
Define, build, and maintain support systems to ensure high availability and performance.
Handle complex cases for the PSO.
Implement automation for system provisioning, self-healing, auto-recovery, deployment, and monitoring.
Perform incident response and root cause analysis (RCA) for critical system failures.
Monitor system performance and establish Service-Level Indicators (SLIs) and Service-Level Objectives (SLOs).
Collaborate with Development and Operations to integrate reliability best practices, including zero-downtime architecture.
Proactively identify and remediate performance issues.
Work closely with Product T&E, ICE, and Service Architects for new product productization as SGS technical expert.
Coordinate with internal and external stakeholders to improve service performance and ensure high availability.
Ensure Operations readiness to support new products.
Accountable within SGS for in-scope product availability and performance.
Problem Management
Conduct thorough problem investigations and root cause analyses to diagnose recurring incidents and service disruptions.
Coordinate with Incident Management teams and collaborate with PSOs and Engineering/Product teams to implement permanent solutions.
Monitor effectiveness of problem resolution activities and provide regular reporting to ensure continuous…