JobRaahGet matched free

Jobs

Lead Site Reliability Engineer

rb · San Francisco, CA · United States · On-site

Posted Oct 5, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

Company Federal Reserve Bank of San Francisco When you join the Federal Reserve—the nation's central bank—you’ll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems. We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we’re building a dynamic and diverse team for our future. The Federal Reserve Financial Services portfolio provides technology capabilities that power the U.S. payment systems infrastructure. This portfolio enables critical payment services that are foundational to the nation's financial system, focusing on delivering speed, resilience, and choice to meet evolving marketplace needs. We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence. Responsibilities System Reliability & Performance •    Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure •    Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance •    Lead incident response, conduct root cause analysis, and implement preventive measures •    Develop and maintain disaster recovery and business continuity plans Infrastructure & Automation •    Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform) •    Automate deployment pipelines, monitoring, and operational workflows •    Optimize cloud resource utilization and cost management Engineering & Development •    Build and maintain internal tools and services to improve operational efficiency •    Collaborate with development teams to implement reliability best practices •    Conduct code reviews and provide technical guidance on system design •    Develop monitoring solutions, alerting systems, and observability frameworks Security & Compliance •    Integrate security practices into CI/CD pipelines (SAST/DAST) •    Implement and maintain security controls across infrastructure and applications •    Ensure compliance with industry standards and regulatory requirements •    Conduct security assessments and vulnerability management Leadership & Collaboration •    Mentor junior SRE team members and promote SRE culture across the organization •    Partner with software engineering teams to improve system reliability •    Drive technical initiatives and contribute to architectural decisions •    Document processes, runbooks, and technical specifications   Software Engineering: Strong proficiency in  Java ,  Python , and  Node.js Experience with microservices architecture and distributed systems Solid understanding of data…