JobRaahGet matched free

Jobs

Software Engineer (Infrastructure)

Thundercompute · San Francisco · United States · On-site

Pay: USD 150,000 – 250,000 a year

Posted Jul 21, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

COMPANY The world is building massive amounts of GPU capacity. Meanwhile, deployed GPUs are only 20% utilized. This is because GPUs are not virtualized, while every other type of hardware is. For example CPUs and storage are allocated through virtual abstractions which efficiently manage the physical hardware, while GPUs are statically allocated on a one-to-one basis. Thunder Compute is building this virtualization layer for GPUs. We have raised over $17.5M from Matrix Partners, Y Combinator, and leading angels from Coreweave, Microsoft, Cognition, and Anthropic. Leading solutions for underutilization sit at the workload layer and are therefore only able to optimize specific use cases. We believe the ideal cluster optimization solution must be invisible to developers and compatible with all workloads; hence, it must sit at the systems layer. We are a team of systems researchers productionizing cutting-edge GPU virtualization research to build this general-purpose optimization layer. Concretely, our virtualization library abstracts GPUs across TCP networking. We use a userspace shim library, loaded through LD_PRELOAD, to intercept CUDA calls and send them over gRPC to a host server connected to a physical GPU elsewhere in the data center. This enables something like “Ceph for GPUs”: GPUs become network resources that can be abstracted, pooled, and dynamically allocated across a cluster to improve utilization without requiring developers to modify their workloads. ROLE Your work will focus on building the cloud infrastructure surrounding our GPU virtualization layer. This includes the Go backbone of our cloud platform, Kubernetes-based orchestration, production reliability, networking, storage, billing infrastructure, and the systems used to deploy and operate GPU capacity at scale. You will take ownership of complex infrastructure from early design through production deployment. Example projects may include: - Building control-plane services for provisioning and managing virtual GPU instances - Designing reliable systems for GPU allocation, scheduling, and lifecycle management - Improving our unconventional Kubernetes deployment, which acts as a form of hypervisor for customer workloads - Building infrastructure for networking, storage, authentication, billing, and usage metering - Automating the deployment and operation of GPU hosts across cloud providers and customer data centers - Debugging failures across customer workloads, Kubernetes, our control plane, the network, and physical GPU infrastructure - Designing systems for failure recovery, capacity management, observability, and incident response - Improving the security, reliability, and operational simplicity of the platform as it scales - Working directly with customers to diagnose problems and deploy Thunder Compute in new environments You will spend your days bouncing between the weeds of complicated production infrastructure that is live and used by…