JobRaahGet matched free

Jobs

Senior AI Scheduling & Orchestration Engineer

Bitdeer Technologies Group · Singapore, SG · On-site

Posted Sep 16, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About Bitdeer: Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence. Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia. What you will be responsible for: Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling. Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic. Leverage Kubernetes Dynamic Resource Allocation (DRA) and custom scheduler plugins to manage complex accelerator requests natively. Architect topology-aware pod placement strategies that optimize for low-latency communication via NVLink and InfiniBand fabrics. Implement automated GPU sharing technologies (e.g., MIG, time-slicing) and multi-tenancy isolation policies to maximize cluster-wide utilization. Collaborate with the GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns. Drive the reliability and scalability of the scheduling stack, resolving resource contention and deadlock scenarios in large-scale HPC environments. Mentor junior engineers and conduct design reviews to maintain architectural excellence in our orchestration layer. How you will stand out: Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field. 6+ years of distributed systems engineering, with deep, hands-on expertise in Kubernetes scheduling frameworks and orchestrators. Extensive experience with AI workload execution patterns and distributed training frameworks (e.g., PyTorch Distributed, Ray, MPI). Proven track record of operating, debugging, and scaling scheduling stacks in high-performance computing (HPC) or large-scale production cloud environments. Strong knowledge of GPU hardware architectures and the specific scheduling challenges related to distributed AI training and inference. Experience with infrastructure automation and infrastructure-as-code (e.g., Terraform, Go-based Operators). Excellent technical communication and leadership skills; ability to influence cross-functional teams and align architectural goals. Ability to work in a high-velocity engineering environment and…