Kubernetes Support Engineer (L2/L3)
Rackbank · Indore · India · On-site
Posted Oct 9, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Location: Indore
Employment Type: Full time
About RackBank Datacenters:
RackBank Datacenters is a fast-growing and trusted data center company committed to powering digital transformation with secure, scalable, and efficient IT infrastructure solutions. With state-of-the-art data centers in India and a strong client base across multiple industries, RackBank is focused on delivering next-gen cloud and colocation services.
Our Vision:
To build India’s most reliable and trusted data center and digital infrastructure ecosystem, empowering enterprises and governments with secure, scalable, and future-ready infrastructure.
About the role:
NeevCloud runs production Kubernetes clusters for its GPU cloud platform, with Ceph as the storage backend and a mixed fleet of NVIDIA datacenter and workstation - class GPUs. You will own the day - to - day health of these clusters and their nodes: Keeping them stable, patched and monitored, and fixing issues before customers notice them. As the platform grows, you will also deploy Kubernetes on new GPU Clusters.
What you will do:
Cluster Operations:
Run daily health checks on the production clusters: nodes, control plane, etcd, certificates and capacity.
Plan and carry out Kubernetes version upgrades with a rollback plan and minimal downtime.
Troubleshoot pod failures, evictions (disk and memory pressure), scheduling problems, and networking and DNS issues, including customer pods that carry a direct public IP.
Support VM workloads running on KubeVirt alongside containers.
Manage persistent storage through Ceph CSI (RBD and CephFS): PVC issues, capacity and performance.
Watch Ceph cluster health with the storage team, and act on warnings that affect Kubernetes volumes (slow requests, nearly full OSDs, degraded placement groups).
Maintain monitoring and alerts in SigNoz and Zabbix, and respond to alerts.
Take etcd backups, test restores, and keep disaster recovery steps documented.
Manage RBAC, namespaces, resource quotas and limits, and work with the security team on vulnerability fixes.
Node Maintenance:
Handle the node lifecycle: add and remove worker nodes, and cordon, drain and return nodes after maintenance.
Patch the OS and kernel, and upgrade GPU drivers and firmware in a planned window, with the data centre team where needed.
Use out-of-band access (BMC, IPMI or Redfish) for remote console, power actions and hardware logs.
Find hardware faults (disk, memory, NIC, GPU), raise vendor tickets, and work with the data centre team on replacements.
Find GPU faults using nvidia-smi, DCGM metrics and Xid errors, and decide whether a node needs draining, a driver fix or a hardware ticket.
Check every node before it returns to production: GPU health, driver and firmware versions, network and storage access.
New GPU Cluster builds:
Deploy Kubernetes on new GPU node clusters, from bootstrap to production handover.
Set up the NVIDIA GPU Operator and…