Lead Site Reliability Engineer
Glint Tech Solutions LLC · Buffalo, New York, United States · On-site
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Job Title: Lead Site Reliability Engineer
Location: Remote within the USA, or onsite in Buffalo, NY / Wilmington, DE (client preference for candidates near these areas). New hires are required to work onsite at the client's office for the first 2–3 weeks (treated as a business trip; travel expenses covered by the company).
Company Overview
Glint Tech Solutions is a women-owned, global IT staffing and recruiting firm serving enterprise clients across the USA and Canada.
Project Description
A leading financial services client is seeking a Lead Site Reliability Engineer responsible at the expert level for ensuring the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. This senior individual contributor will design, implement, and improve SRE practices across the software development lifecycle, working closely with application development, infrastructure, platform engineering, and business teams to enhance system resiliency through automation, observability, testing, and proactive operational management, while coaching and influencing others.
Key Responsibilities
Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise SRE best practices
Define, implement, and monitor SLOs, SLIs, and error budgets for critical business services
Develop observability strategies using Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting
Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints
Lead incident response for high-severity production events and facilitate Root Cause Analysis (RCA)
Drive operational excellence through automation of deployments, recovery procedures, and reliability controls
Design and execute automated regression testing strategies to validate stability and performance
Create and maintain Infrastructure as Code (IaC) solutions using Terraform
Support and optimize Microsoft Azure environments, including App Services, scaling, and deployment automation
Utilize Azure Monitor, Application Insights, and Log Analytics to improve platform visibility
Drive performance testing, resiliency testing, and disaster recovery preparedness
Lead capacity planning, performance tuning, and workload optimization
Develop operational runbooks, incident playbooks, and standard operating procedures
Mentor engineers on observability, cloud engineering, automation, and SRE principles
Adhere to Company risk and regulatory standards, policies, and controls
Mandatory Skills
Strong hands-on experience with Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics collection, and centralized logging
Proven experience designing and executing automated regression testing frameworks
Strong proficiency in Infrastructure as Code (IaC) using Terraform
Experience with CI/CD…