JobRaahGet matched free

Jobs

Senior Software Engineer

Turing · São Paulo, Brazil · On-site

Posted Sep 28, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About Turing Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models. Learn more at www.turing.com . ABOUT THE ROLE Turing is the world's leading research accelerator for frontier AI labs. The STEM Horizontal team produces the high-signal STEM data those labs train on, and this role builds the machinery behind it: the task-authoring tools, data pipelines, evaluation harnesses, sandboxed execution environments, and agentic workflows our researchers and domain experts work inside every day. This sits at the seam between engineering and research. You'll own production systems end-to-end, and you'll be the person a researcher pulls in to ask why an eval is producing misleading numbers. Correctness and reproducibility are the product: a subtle bug here doesn't crash anything, it quietly poisons a training run and costs weeks. If you want to build systems that visibly shape how frontier models learn to reason about code and science, this is the seat. WHAT YOU'LL DO Own services end-to-end: design, implementation, rollout, and production operation. Build backend services and the data pipelines that generate, validate, score, and version datasets, reproducible from a commit and a config. Ship data-dense internal interfaces: virtualised tables, server-side filtering, review and annotation UIs that stay fast at tens of thousands of rows. Build and operate agent harnesses , multi-step loops with tool use, retries, and structured output, running against real repos and test suites. Build sandboxed execution environments where model-generated code runs safely and deterministically at volume, and stand up the RL environments and eval harnesses on top. Trace and debug agent runs end-to-end: where the loop stalled, which tool call failed, why the grader disagreed with the human. Own infrastructure as code and CI/CD; debug production issues across the stack and drive the reliability work that follows. Raise the bar through code review, design docs, and innovation, partnering with researchers and quality owners on what “good” means. TECH STACK Backend Python · FastAPI · Django / Django Ninja · PostgreSQL · MongoDB · REST +…