Principal Site Reliability Engineer (Kubernetes Required) - Hybrid
factset · Norwalk, CT, USA · United States · Hybrid
Pay: USD 190,000 – 220,000 a year
Posted Sep 24, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
FactSet creates flexible, open data and software solutions for over 200,000 investment professionals worldwide, providing instant access to financial data and analytics that investors use to make crucial decisions.
At FactSet, our values are the foundation of everything we do. They express how we act and operate, serve as a compass in our decision-making, and play a big role in how we treat each other, our clients, and our communities. We believe that the best ideas can come from anyone, anywhere, at any time, and that curiosity is the key to anticipating our clients’ needs and exceeding their expectations.
About the Role
We are looking for a skilled and motivated Principal Site Reliability Engineer to join our team. In this role, you will be responsible for ensuring the reliability, scalability, and performance of our systems and services. You will work closely with development and operations teams to build and maintain robust infrastructure, automate processes, and drive engineering best practices.
Key Responsibilities
Monitor, maintain, and improve the reliability and availability of production systems
Respond to and resolve incidents, conducting thorough post-mortems to prevent recurrence
Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
Collaborate with development teams to build reliability into services from the ground up
Design and implement automation to reduce toil and improve operational efficiency
Participate in an on-call rotation to support critical systems
Contribute to capacity planning and performance optimization efforts
Document systems, processes, and runbooks to support the wider team
Minimum Experience:
8+ years’ experience ensuring the reliability, scalability, and performance of our systems and services
Required Technical Skills
Kubernetes (Required)
Hands-on experience deploying, managing, and troubleshooting workloads in Kubernetes
Strong understanding of core Kubernetes concepts including Pods, Deployments, Services, ConfigMaps, and Ingress
Experience with Kubernetes cluster management and administration
Familiarity with Helm for application packaging and deployment
Understanding of Kubernetes networking, storage, and security best practices
Additional Technical Skills
Cloud Platforms: (e.g. AWS, GCP, Azure)
CI/CD Tooling: (e.g. GitHub Actions, ArgoCD, Harness)
Monitoring & Observability: (e.g. Prometheus, Grafana, Coralogix, OpenTelemetry)
Infrastructure as Code: (e.g. Terraform, Pulumi)
Config Management : (e.g. Ansible, Puppet, Chef)
Programming/Scripting: (e.g. Python, Go, Bash)
Soft Skills & General Requirements
Strong problem-solving and analytical skills with a methodical approach to troubleshooting
Excellent communication skills with the ability to collaborate across technical and non-technical teams
A proactive mindset with a focus on automation and continuous improvement
Ability to work…