Site Reliability Engineer
Tvsnext · Chennai · India · Hybrid
Posted Sep 7, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
We are looking for Site Reliability Engineer - Chennai
About the job
What You ‘ll Do
You will join our high-performance Cloud Engineering team and play a critical role in designing, automating, and maintaining highly available, scalable, and resilient cloud infrastructure platforms that power enterprise-grade applications and digital experiences.
Design, implement, and manage scalable AWS cloud infrastructure to support mission-critical business applications.
Build, deploy, and maintain containerized workloads using AWS ECS and AWS Fargate.
Manage and optimize relational database infrastructure leveraging AWS RDS and PostgreSQL.
Develop and maintain Infrastructure as Code (IaC) frameworks to enable automated, repeatable, and consistent infrastructure deployments.
Design, implement, and optimize CI/CD pipelines to improve deployment efficiency, reliability, and release velocity.
Implement and optimize high-performance caching strategies to improve application responsiveness, scalability, and overall system performance.
Ensure infrastructure availability, reliability, scalability, and security across production and non-production environments.
Monitor platform health, system performance, and application availability while proactively identifying and addressing potential issues.
Troubleshoot and resolve complex infrastructure, networking, and platform-related issues across distributed systems.
Perform root cause analysis for production incidents and drive preventive actions to improve system reliability.
Collaborate closely with software engineering, architecture, DevOps, and QA teams to support cloud-native application delivery.
Establish and implement reliability engineering best practices including observability, automation, incident management, and capacity planning.
Document operational procedures, architecture decisions, and best practices to ensure knowledge sharing and operational excellence.
Contribute to continuous improvement initiatives focused on system reliability, operational efficiency, and infrastructure modernization.
What We Seek In You
5+ years of experience designing, implementing, managing, and troubleshooting scalable cloud infrastructure environments.
Strong hands-on expertise in AWS cloud services and cloud-native architecture patterns.
Proven experience working with containerized environments using AWS ECS and AWS Fargate.
Strong experience managing relational database platforms including AWS RDS and PostgreSQL.
Advanced proficiency in Infrastructure as Code (IaC) tools such as Terraform, AWS CloudFormation, or equivalent technologies.
Extensive experience designing and implementing CI/CD automation pipelines using modern DevOps practices and tools.
Strong expertise implementing and optimizing high-performance caching layers and distributed caching solutions.
Solid understanding of cloud architecture principles including high availability, scalability, fault tolerance, and disaster…