Site Reliability Engineer
Alvaria · TX, US · United States · On-site
Posted Oct 8, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
About Aspect Software
Building on more than 50 years of industry experience, Aspect Software is reimagining workforce management through cloud technology, AI, automation, and human-centered innovation. Our Workforce Engagement Management solutions help organizations solve complex workforce challenges, improve operational performance, and deliver better employee and customer experiences.
We foster a collaborative environment where engineers work across teams and geographies to build secure, scalable, and dependable software. Join us as we modernize intelligent workforce systems used by organizations around the world.
Position Overview
We are seeking a Site Reliability Engineer to help build and operate reliable, secure, and efficient cloud services. In this role, you will combine software engineering and systems expertise to improve availability, performance, scalability, observability, and operational readiness across our SaaS platforms.
This person will partner closely with application engineering, platform, security, and product teams to define service-level objectives, automate repetitive work, strengthen incident response, and design systems that remain resilient as they scale. The ideal candidate is a pragmatic problem-solver who measures what matters, learns from failures, and improves systems through automation rather than manual intervention.
Key Responsibilities
Design, build, and operate highly available, scalable, and secure production services and cloud infrastructure
Partner with engineering teams to define and maintain service-level indicators (SLIs), service-level objectives (SLOs), error budgets, and actionable alerts
Build automation and self-service tooling that reduces toil, accelerates safe delivery, and improves operational consistency
Improve observability through meaningful metrics, logs, distributed tracing, dashboards, synthetic checks, and health monitoring
Participate in an on-call rotation and lead or support incident response, communication, mitigation, and recovery
Facilitate blameless post-incident reviews and ensure corrective actions address systemic causes
Improve CI/CD pipelines using automated validation, deployment health gates, progressive delivery, canary releases, and reliable rollback strategies
Review system designs for reliability, performance, capacity, security, recoverability, and cost efficiency
Strengthen resilience through redundancy, autoscaling, fault-tolerant patterns, disaster recovery planning, and appropriate load or chaos testing
Troubleshoot complex issues across applications, infrastructure, networking, databases, and third-party dependencies
Develop and maintain runbooks, operational standards, architectural documentation, and support procedures
Collaborate across teams and time zones while clearly communicating risk, trade-offs, incident status, and reliability priorities
Required Qualifications
Bachelor’s degree in Computer Science, Engineering, or a related…