AI Builder — Working Student
EggAI · Tübingen, Germany · On-site
Posted Sep 8, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
EggAI Labs · Working Student · In Person (Tübingen, Germany)
Designing lifecycle evaluation for nondeterministic domain-specific AI systems.
EggAI Labs
We're opening the EggAI Labs in Tübingen. The Labs purpose is to test new AI capabilities, identify what works, how they'd improve delivery, and then build assets to support teams and clients. EggAI's mission cannot be achieved through one-off client projects alone and the Labs is the way to adapt and scale.
As part of the first cohort, you will join a small peer group and work directly with EggAI's CAIO, a Tübingen alumnus. You will also collaborate with our engineering, product, and project leads, who will bring you insights from client problems, help you understand their context and challenge your solutions for them. Your fellow AI Builders will be deliberately chosen to span different passions and strenghts. You will have room to explore and be expected to turn that exploration into something useful. This is not a client-facing role.
The problem: EvalOps
How do we prove an AI system acts as expected under various conditions?
We will explore EvalOps : the methods and tools used to check whether domain-specific AI systems meet their behavioral requirements. Among the key challenges for operating AI systems are nondeterminism, ground-truth curation, evaluation metric-design, latency and costs. EvalOps need to overcome all of them in a disciplined way.
These questions will guide the initial work:
Regression suites : How can domain experts turn traces, code, and prompts into useful evals in a way that scales?
Coverage estimation : How can we quantify the gaps between an eval suite and the range of expected agent behaviors?
Live monitoring : Is it possible to detect anomalous behaviour in production with a similar precision to regression evals?
Eval analysis : How much evidence is enough to act (e.g., deploy, debug, roll back a feature, choose a different model)? You will help refine them, test approaches, create assets, and pose new questions.
The role
As an AI Builder, you will take an open EvalOps problem from investigation to a working prototype, ready to be tested. Whatever your focus, you will be expected to take a problem through to a useful result. Here is what that means in practice:
Research emerging AI and evaluation capabilities and distinguish substance from hype.
Understand the user, workflow, domain, and organisational problem around the work.
Evaluate usefulness, quality, safety, and reliability with appropriate evidence.
Communicate decisions and learning clearly.
Create assets that other people can understand, use, and improve.
Code and build end to end.
This shared foundation leaves room for different strengths and interests. Each AI Builder will then develop greater depth in one of three pillars.
Three pillars, different passions
Pillar
Driving question
Possible EvalOps focus
Product
What should we build?
Workflows and…