Senior Site Reliability Engineer, Events
hireworks.io · Rio de Janeiro, State of Rio de Janeiro, Brazil · On-site
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
About hireworks
hireworks is building a community of top talent in key international markets by unlocking unparalleled access to positions at leading U.S. based companies. As your employer, hireworks will ensure you have a seamless interview, onboarding, and employee experience - providing ongoing support and resources along the way. Established in 2023, hireworks is forging corp-to-corp relationships with leading U.S. based organizations looking to grow their teams with best-in-class talent around the world. Working with hireworks means unlocking access to a network of local peers and mentors and career opportunities through our client network.
About Our client
A leading provider of innovative software and services to K-12 schools and educational institutions worldwide. Their mission is to empower schools to transform education through technology. They believe in the power of education to shape the future, and are committed to helping schools prepare students for success and make a positive impact on the world.
Position Overview
Our client's Events platform powers event marketing, ticket sales, ticket management, and on-site ticket distribution and scanning for organizers running live events. It's a high-throughput, event-driven application on AWS, and it needs a Senior SRE to keep it reliable, well-instrumented, and continuously improving in place.
You will operate and strengthen the platform's AWS infrastructure, CI/CD, and observability, working closely with the engineers who build and evolve the application so the system holds up under real event-day load — on-sales, high-traffic scans, and everything in between.
Key Responsibilities
Operate and maintain the Events platform's AWS infrastructure — compute, data stores, CDN, and the event bus connecting bookings, tickets, users, and scanning
Build and extend observability — dashboards, tracing, and alerting that catches problems before customers do, especially around live on-sales and event days
Maintain and improve CI/CD pipelines for the booking, box office, and ticket-scanning applications, working toward deployments that are repeatable, auditable, and low-risk
Partner with the engineers building and evolving the Events platform, understanding the event-driven architecture well enough to contribute code, diagnose production issues, and suggest infrastructure improvements
Participate in on-call and incident response for the platform, with particular attention to the operational patterns unique to ticket scanning at scale
Apply IaC discipline to infrastructure changes, avoiding manual, untracked changes to production
Use AI-assisted development tools (Claude Code or comparable) as part of your workflow to navigate and document a complex codebase and infrastructure estate
Required Qualifications
~6 years of experience as an SRE, DevOps engineer, or backend/infrastructure engineer supporting production systems
Strong, hands-on AWS experience — comfortable…