JobRaahGet matched free

Jobs

NOC / SRE Lead – NOC Operations

Phykon · Thiruvananthapuram · India · On-site

Posted Aug 12, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

Location: Thiruvananthapuram, Kerala, India (Onsite) Experience: 8+ Years Employment Type: Full-Time Department: Network Operations Center (NOC) Reports To: [Head of Operations / Engineering – to be confirmed] About the Role We are seeking an experienced NOC / SRE Lead to own the health, observability, and incident response for a connected fleet of field-deployed systems and the data center infrastructure that supports it. This is the keystone L3 role in our Network Operations Center (NOC). You will own deep diagnosis and root-cause analysis on the running system, command the hardest escalations, and build the standards, runbooks, observability architecture, and operator-certification program that the rest of the NOC inherits. As the L3 Lead, you will act as the escalation point above L2 support teams, taking ownership of the most complex incidents while keeping escalations to L4 product engineering limited to genuine code and firmware issues — protecting both fleet uptime and engineering velocity. This is a greenfield opportunity. In the near term, the role combines a hands-on L3 technical function with NOC leadership. As the NOC scales, the role is designed to grow into a Principal SRE or NOC Manager track. Key Responsibilities 1. Fleet Health & Observability Architect the monitoring stack the NOC runs on — scrape architecture, alert rules, dashboards, and SLOs — rather than only consuming it. Own end-to-end visibility of the connected fleet across data center, edge Kubernetes, Linux, and network layers. Continuously tune alerting to reduce noise and catch degradation early. 2. Data Center & Infrastructure Operations Oversee the health and availability of data center and colocation infrastructure — compute, storage, network, power, and cooling dependencies — supporting the fleet. Coordinate with data center operators, colocation providers, and remote-hands teams for physical interventions, maintenance windows, and capacity changes. Manage hardware fault handling, RMA workflows, and site-level incident response across distributed data center and edge sites. Coordinate with telecommunications providers, colocation partners, cloud vendors, and infrastructure service providers to resolve service-impacting issues and maintain operational continuity. 3. Incident Management & Response Serve as incident commander on the hardest escalations, including Sev1/P1 incidents, and drive them to resolution within SLA. Perform deep diagnosis and root-cause analysis on the running system using logs, metrics, and telemetry. Diagnose complex issues across: Network connectivity, routing, NAT/CGNAT, and tunnels Degraded or intermittent links to field-deployed hardware Production Kubernetes and Linux platform dependencies at the edge Assume end-to-end ownership of major incidents, service disruptions, and customer escalations until resolution and formal closure. Act as the highest operational escalation point within the NOC for…