JobRaahGet matched free

Jobs

Site Reliability Engineer

Fabric · New York, United States · Remote

Pay: USD 135,000 – 160,000 a year

Posted Aug 31, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About Fabric Health At Fabric Health, we are powering boundless care by solving healthcare’s biggest challenge: clinical capacity. We aren’t here to disrupt healthcare; we’re here to fix it. We unify the care journey from intake to treatment, using intelligent automation to remove administrative burdens and make care delivery 2-10x more efficient. Our technology empowers clinicians to move faster and focus on what matters most: the patient. We are a mission-driven team of brilliant minds trusted by leading organizations including Intermountain Health, OSF HealthCare, SSM Health, and MUSC Health. Our vision is backed by premier investors such as Thrive Capital, GV (Google Ventures), General Catalyst, and Salesforce Ventures. We move quickly for good reason, listen deeply to solve big challenges, and build products with the same care and quality we’d want for our own loved ones. Learn more: About Us | News & Press | LinkedIn | Careers About the Role We are looking for a Site Reliability Engineer to help us evolve and safeguard the infrastructure powering healthcare experiences for millions of patients, all while keeping operational toil to a minimum for the Fabric tech community. What You'll Do As a Site Reliability Engineer, you will act as an architect of our AWS and Kubernetes (EKS) platform, playing an active part in shaping the strategy to make it resilient, scalable, and compliant. Your primary responsibilities include: Infrastructure & Kubernetes Orchestration Designing, deploying, and maintaining Kubernetes (EKS) clusters for enterprise-grade availability. Optimizing the footprint of infrastructure objects across core AWS services (EC2, RDS, S3) for performance, cost, and reliability. Evolving a scalable infrastructure management platform, with the right interfaces and guardrails to maximize engineering agency at minimal cognitive load. Defining golden paths for workload orchestration and integration with infrastructure dependencies, ensuring the right way to do things is also the easiest. Automation & AI-Assisted Operations Providing robust, reusable GitHub Actions components to streamline the software delivery lifecycle. Developing internal tools that replace manual operations with intelligent, autonomous systems. Exploring and deploying agentic workflows for AI-assisted runbooks that automate complex, error-prone procedures and repetitive tasks. Observability & Incident Management Driving the evolution of observability practices by maintaining reliable mechanisms to collect the metrics, traces, and logs needed to meet SLOs. Leading incident response efforts and facilitating the blameless postmortems that help systematically reduce recovery time (MTTR). Defining and monitoring platform-level SLIs and SLOs to ensure it consistently meets rigorous healthcare performance standards. Compliance & Collaboration Ensuring every piece of infrastructure is continuously compliant with HIPAA and other critical…