JobRaahGet matched free

Jobs

Senior Site Reliability Engineer - FedRAMP

ffive · Reston · On-site

Pay: USD 161,900 – 242,900 a year

Posted Oct 7, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation.    Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive. Role Summary We are seeking an experienced, security-focused Senior Site Reliability Engineer (Senior SRE) to drive the reliability, architectural design, and continuous compliance of our FedRAMP-authorized cloud platform. In this senior role, you will combine hands-on operational leadership with infrastructure architecture, technical governance, and audit readiness across AWS (Commercial and GovCloud) and Kubernetes environments. As a Senior SRE, you will design and operate highly resilient, multi-cluster Amazon EKS infrastructure, direct observability architectures using Prometheus and Grafana, maintain automated deployment pipelines using GitLab CI/CD, and lead technical response in a 24/7 rotational on-call schedule. You will also serve as a key technical liaison for Third-Party Assessment Organization (3PAO) FedRAMP audits, ensuring strict adherence to NIST SP 800-53 controls, architectural security patterns, and vulnerability management SLAs. Role Snapshot Primary Focus: Production reliability, hybrid cloud & edge infrastructure, FedRAMP Compliance, Secure Architectural Design & L3 Escalation Support. Cloud & Edge Platforms: AWS (GovCloud / Commercial EKS), EKS, S3, KMS, RDS, IAM bare-metal on-premises edge servers. Core Tooling: Kubernetes, ArgoCD, Terraform, GitLab CI/CD (FIPS-compliant runners). Observability & Logging: Prometheus, Grafana, Alertmanager, Slack, Elasticsearch, Kibana, SIEM integration. Programming: Python, Go (Golang), Bash. On-Call Rotation: L3 escalation support Key Responsibilities 1. L3 On-Call Escalation & Production Support • Act as the final technical escalation tier (L3 support) for critical platform and production incidents, troubleshooting complex issues escalated by L1/L2 operations or customer support teams. • Lead high-priority incident response bridges, coordinating across development, security, and networking teams to drive rapid issue containment and resolution under strict service-level agreements (SLAs). • Troubleshoot deep, transient infrastructure errors (e.g., routing, network package drops, Kubernetes control-plane failures) that go beyond standard SOPs. • Establish clear, documented escalation pathways, translating complex L3 resolutions into actionable runbooks and empower L1/L2 teams 2. FedRAMP Compliance, Audit Readiness & Log Auditing • Implement and enforce security controls aligned with the NIST SP 800-53 framework to achieve…