Machine Learning Engineer - Infrastructure
Transparent Search Group · San Francisco, California, United States · On-site
Pay: USD 200,000 – 400,000 a year
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Machine Learning Engineer - Infrastructure
Company: Causal Labs
Location: San Francisco, CA (South Park office, in person 5 days per week; relocation provided)
Compensation: $200,000 - $400,000 + highly competitive early-stage equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers; can sponsor visas
About Causal Labs
Causal Labs is pursuing general causal intelligence: AI that can predict the future and identify the actions that change it. It is building a Large Physics foundation Model (LPM), because domains governed by physics have inherent cause-and-effect structure that visual or textual data lacks. Its starting domain is weather, the most observed physical system on earth, with rapid ground-truth feedback and data volumes that dwarf LLM training sets.
The founders come from Cruise, Google Research and Meta. The company is about 10 people in San Francisco, growing to around 35 this year, and is backed by Kindred Ventures, Refactor and BoxGroup.
The Role
Causal Labs is hiring infrastructure engineers to tackle the unsolved training and inference challenges of a Large Physics foundation Model. The work demands deep expertise in standing up distributed training clusters and optimizing performance for large models. If you have built large-scale ML infrastructure for language, vision, robotics or biology models and want to bet on a counterintuitive technical thesis, this is the role.
What You Will Do
Design, deploy and maintain large distributed ML training and inference clusters.
Build efficient, scalable end-to-end pipelines for petabyte-scale datasets and model training across the ML lifecycle.
Research and test training approaches, including parallelization techniques and numerical-precision trade-offs across model scales.
Analyze, profile and debug low-level GPU operations to optimize performance.
Bring new ideas from current research into the stack.
What You Bring
2-10 years building large-scale ML infrastructure for core foundation models (not fine-tuning)
Deep expertise optimizing large-scale training and inference workloads
Proficiency with distributed training frameworks (FSDP, DeepSpeed)
Knowledge of cloud platforms (GCP, AWS or Azure) and containers/orchestration (Kubernetes, Docker)
Experience with distributed task management and scalable model serving architectures
Strong grasp of monitoring, logging and observability for ML systems
Ability to work in person in San Francisco 5 days a week
Nice to Have
Experience at a science or physical AI company (self-driving, robotics, biology, climate/weather)
Generalist experience across the ML lifecycle
Low-level GPU performance optimization and debugging (CUDA, JAX)
Interview Process
Initial call (30 min), technical screen, onsite day.
Tech Stack
FSDP, DeepSpeed, NVIDIA GPUs, Python, C++, Linux, Kubernetes, Docker, GCP/AWS/Azure