JobRaahGet matched free

Jobs

Senior AI Storage Infrastructure Engineer

Bitdeer Technologies Group · Singapore, SG · On-site

Posted Jul 28, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About Bitdeer: Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence. Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia. What you will be responsible for: Design, deploy, and maintain robust Container Storage Interface (CSI) drivers for high-performance parallel file systems (e.g., Weka, Lustre, DAOS, VAST). Architect and implement GPUDirect Storage (GDS) integrations to enable direct memory access (DMA) between NVMe drives and GPU memory, bypassing CPU bottlenecks. Develop and manage local NVMe caching strategies for rapid, low-latency loading of massive model weights and datasets during distributed training. Optimize IOPS, throughput, and latency profiles across the entire containerized storage stack, from the storage array to the container runtime. Collaborate with the GPU Systems & Fabric team to ensure the storage layer is fully optimized for RDMA and high-speed interconnects (InfiniBand, RoCE). Implement automated monitoring and alerting for storage performance, detecting and mitigating I/O contention or hardware degradation before it impacts production jobs. Define storage policies, quota management, and multi-tenancy isolation strategies within Kubernetes to ensure fair resource sharing for customer workloads. Mentor junior engineers and drive architectural design reviews to maintain high standards of reliability and performance across the infrastructure team. How you will stand out: Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field. 5+ years of experience in distributed storage systems and high-performance file systems, with a deep understanding of POSIX compliance and file I/O semantics. Deep expertise in the Kubernetes CSI paradigm, including building or extending volume plugins and storage operators. Strong hands-on experience with block/file I/O at the Linux OS level and kernel-level performance tuning. Familiarity with high-throughput networking protocols (RDMA, InfiniBand, RoCE) and how they interact with storage subsystems. Proven track record of operating, debugging, and scaling large-scale storage environments in production or HPC settings. Experience with infrastructure automation tools (e.g., Terraform, Ansible) and CI/CD pipelines Excellent technical…