JobRaahGet matched free

Jobs

Staff Slurm Cluster & HPC Engineer

Bitdeer Technologies Group · San Jose, CA · United States · Remote

Posted Aug 25, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia. To learn more, visit https://ir.bitdeer.com/ Position Overview We are seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm as a first-class, productized scheduling layer across that fleet. This person is the single technical owner of Slurm cluster architecture, multi-tenant scheduling policy, and cluster reliability on both bare-metal and VM-based GPU nodes, and will lead our adoption of the Slinky operator stack (slurm-operator, and slurm-bridge where it fits) so that Slurm and Kubernetes workloads can share the same GPU pool. The role is deeply hands-on, customer-facing during onboarding and escalations, and sets the engineering standard the rest of the platform team builds on. Key Responsibilities Slurm cluster architecture and lifecycle — Design, deploy, and operate production Slurm clusters on bare metal and VMs: slurmctld/slurmdbd high availability, slurmrestd, configless slurmd, SACK/MUNGE and JWT authentication, and rolling version upgrades on live clusters without losing running jobs. Topology-aware scheduling for GPU fabrics — Model the physical fabric in topology.conf — topology/tree for rail-optimized InfiniBand/RoCE designs and topology/block for NVLink domains such as GB200/GB300 NVL72 — and prove placement quality with NCCL bandwidth and multi-node training validation rather than assumption. Multi-tenant scheduling policy — Own the account/association tree, partitions, QOS, fairshare, preemption, reservations, and per-tenant TRES limits. Enforce fail-closed defaults: an unresolved tenant identity or an empty entitlement set must deny, never degrade into unrestricted access. Slinky on Kubernetes — Lead implementation of the Slinky slurm-operator, including its NodeSet, LoginSet, Accounting, RestAPI, and Token custom resources, cert-manager and Helm-based delivery, shared parallel-storage mounts, and login pods running sackd/sshd. Evaluate and pilot slurm-bridge for co-scheduling Kubernetes Pods, PodGroups, Jobs, JobSets, and LeaderWorkerSets through the Slurm scheduler, and document its constraints — notably exclusive whole-node allocation — before any customer exposure. Elastic capacity between Slurm and Kubernetes — Use Slurm cloud and power-save mechanisms (ResumeProgram/SuspendProgram, SuspendTime,…