JobRaahGet matched free

Jobs

Cloud Systems Engineer - Site Reliability

TherapyNotes.com · Remote · Philadelphia, Pennsylvania, United States · Remote

Pay: USD 110,000 – 150,000 a year

Posted Sep 28, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About Us TherapyNotes is the go-to superhero for behavioral health Practice Management and EHR software! Our top-notch SaaS solution handles scheduling, billing, documenting, telehealth, and more so clinicians can focus on awesome patient care. We're a dynamic team of pros who love to innovate and push the envelope, keeping our software cutting-edge. Join us, and let's revolutionize behavioral health software together while making a real difference! About The Position We are seeking a Site Reliability Engineer to improve the reliability and operability of the production services and shared platforms supporting our growing 24×7 SaaS environment . In this role, you will apply software and systems engineering practices to improve availability, performance, scalability, resilience, observability, incident response, and operational automation . You will partner with software development, infrastructure, database, security, and other technology teams to establish measurable reliability goals, reduce operational toil, and ensure services are supportable throughout their lifecycle. If you are passionate about building reliable systems, solving complex production problems, and driving continuous improvement, we want to hear from you. Requirements BS degree in Information Systems, Engineering, or equivalent experience. 5+ years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or SRE. Experience designing and operating production systems using cloud-based compute, storage, networking, and containerization technologies; Azure and Kubernetes preferred . Strong Linux systems and networking fundamentals, with experience troubleshooting complex distributed systems in production. Expertise with an observability platform; Datadog experience strongly preferred . Experience with Prometheus, Grafana, New Relic, or equivalent platforms is also valuable. Experience with scripting and operational automation using tools such as Bash, PowerShell, or Python, along with infrastructure-as-code and configuration-management practices. Experience participating in production on-call rotations, incident response, root cause analysis, and post-incident improvement. Experience working in Agile/DevOps environments and operating production services using ITSM practices where applicable. Prior software development experience—or experience investigating application behavior through code, logs, and distributed traces—is a plus Responsibilities Own and continuously improve how we use Datadog to make reliability visible and actionable across metrics, logs, traces, dashboards, monitors, alerts, and service-level views. Design, implement, and maintain high-availability, high-throughput, data- and compute-intensive critical systems supporting a growing 24×7 SaaS platform. Partner with service owners to define and improve reliability through meaningful SLIs, SLOs, error budgets , actionable alerting, and…