JobRaahGet matched free

Jobs

ML Engineer, Infrastructure

Prior Labs · Berlin · Germany · On-site

Posted Aug 4, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

WHO WE ARE Foundation models transformed text and images. Structured data - the largest and most consequential data format in the world - stayed untouched, until now. What LLMs did for language, we're doing for tables. We pioneered tabular foundation models: TabPFN v2 was a Nature https://www.nature.com/articles/s41586-024-08328-6 cover story, has passed 3.5M+ downloads and 7,500+ GitHub stars, and runs in production from detecting lung disease with Oxford Cancer Analytics https://www.oxcan.org/news/prior-labs-and-oxford-cancer-analytics-partner-to-advance-liquid-biopsy-and-clinical-decision-making-in-lung-disease to preventing train failures with Hitachi https://siliconangle.com/2025/12/01/prior-labs-debuts-tabular-ai-foundation-model-scales-10-million-rows/. The hardest problems - millions of rows, real-time inference, entirely new modalities - are still open, and no one else is working on them at this level. We're a small, highly selective team of 40+ https://priorlabs.ai/about with backgrounds from Google, DeepMind, Meta, Apple, Amazon, Jane Street, and CERN, led by Frank Hutter https://www.linkedin.com/in/frank-hutter-9190b24b/, Noah Hollmann https://www.linkedin.com/in/noah-hollmann-668b9010b/, and Sauraj Gambhir https://www.linkedin.com/in/sauraj-g/, and advised by Bernhard Schölkopf and Turing Award winner Yann LeCun. In July 2026, less than 18 months after our €9M pre-seed, we joined SAP https://priorlabs.ai/blog-posts/priorlabs-sap as an independent frontier AI lab - same team, mission, and open-weights models, now backed by more than €1 billion over four years. ABOUT THE ROLE We spend tens of millions per year on GPU compute to train tabular foundation models. That's not a target, it's what we're running today, and it's growing. The person who owns this infrastructure makes decisions worth millions of dollars: cluster architecture, scheduling efficiency, provider strategy, hardware selection. A wrong call costs six figures. Today we run Slurm on GCP across multiple clusters. We're scaling to multi-cluster, multi-provider infrastructure and evaluating new hardware generations as they come online. You own the full stack, from cluster operations and cost optimization to distributed training performance and the tooling layer that keeps researchers moving fast. You work directly with the research team and understand what they're doing well enough to make infrastructure decisions that actually help them. And this isn't a pure support role. We operate an open environment. If you've got the next SOTA tabular architecture up your sleeve, go ahead and train it. What you'll work on: - Own and evolve multi-cluster GPU infrastructure. Slurm on GCP today, multi-provider and new hardware tomorrow. Architecture, scheduling, reliability, cost optimization - Drive GPU utilization and training throughput: profiling, memory optimization, communication bottlenecks, systems-level debugging of distributed training across large runs -…