Senior SRE
TechGrove by Banyan Software · Bengaluru, India · On-site
Posted Jul 13, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
TechGrove is the Centre of Excellence for Banyan Software, based in Chennai, India. It plays a key role in supporting Banyan’s global businesses through technology, security, and software development. TechGrove brings together India’s deep pool of technical talent with Banyan’s long-term approach to growth, creating a trusted, developer-focused environment where people can do their best work.
Job Title: Senior SRE (Site Reliability Engineer) – Modernized Application Operations
Overview
We are seeking a highly experienced and hands-on SRE to own the operational excellence of the modernized SaaS applications produced by the Banyan AI Factory. This is not a role focused on building the factory itself; instead, you will run the reliability of the modernized applications the factory delivers to our Operating Companies (OpCos).
You will join a team that provides 24x7 coverage with rotating on-call responsibilities, serving as Tier 1 Site Reliability Engineering (SRE) for our OpCos’ distributed applications. Day to day this will include: automated deployments, cloud service integration, application performance and availability monitoring/observability, and security incident response across our two target clouds — Amazon Web Services (AWS) and Microsoft Azure. The ideal candidate has a track record of keeping secure, highly available production systems running at scale.
Key Responsibilities
24x7 Operations & On-Call: Operate as part of a team providing round-the-clock coverage of OpCo containerized applications, participating in a rotating on-call schedule to ensure continuous availability and rapid response.
Tier 1 SRE & Operations: Serve as Tier 1 SRE for the modernized applications, managing day-to-day cloud integrations across our two target clouds — AWS OR Azure — to keep production systems healthy, performant, and secure.
Performance & Availability Monitoring/Observability: Implement and maintain robust application observability tooling (monitoring, logging, tracing) to track performance and availability, proactively detect degradation, and drive down mean-time-to-detect and mean-time-to-resolve.
Disaster Recovery and Service Restoration: Develop, maintain, test, and execute disaster recovery and business continuity procedures. Ensure the timely recovery and restoration of services following geographic disruptions, cyber incidents, infrastructure failures, or other disaster events.
Security Incident Response: Respond to security incidents and operational events affecting OpCo SaaS platforms, executing established runbooks, coordinating remediation
Automation & Infrastructure-as-Code : Use Infrastructure-as-Code (Terraform) and CI/CD pipelines (e.g., GitHub Actions, GitLab CI) to manage, deploy, and automate the operational environments of modernized applications, reducing toil and improving consistency.
AI Agents & DevSecOps Scale: Build scale in our DevSecOps practice by designing, building, and operating AI agents that…