Machine Learning / Data Engineer
Turing · Brazil; Colombia, Huila, Colombia; São Paulo, Brazil · Remote
Posted Oct 6, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
About Turing
Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models. Learn more at www.turing.com .
Senior ML & Data Engineer — Data Quality & Sensitive Data Compliance
This is a full-time remote role based in Brazil or Colombia.
About the role
Enterprise data flows through our connectors, gets processed, and passes through a sanitization layer before anything downstream touches it. Two things have to be true at every step: the data is what we think it is, and no sensitive information — PII, PHI, company identifiable information (CII), or financial data — gets through. You'll own both. You'll do this primarily by building the machine learning that detects sensitive entities in text and image data and replaces them consistently at scale.
This is a hands-on IC engineering role with a QA mindset. You'll build the detection models, validation infrastructure, adversarial test sets, and audit processes that let us make strong claims about data quality and de-identification performance — and back them up with evidence. You'll work closely with a senior ML lead, with no client-facing responsibilities.
What you'll do
Data Quality
Run deep dives into enterprise data to assess quality: topic coherence across connectors, domain depth within connectors, completeness, and consistency
Design and automate validation suites for data pipelines — schema checks, completeness, drift detection, and reconciliation across raw → processed → sanitized stages
Surface and characterize quality issues in ways that engineering and product can act on
Sensitive data compliance (PII / PHI / CII / financial)
Design, train, and evaluate ML models (NER and other approaches) that detect sensitive entities across text and image-based documents such as scans, invoices, and presentations
Build replacement pipelines that substitute detected entities with coherent alternatives, so the same entity always maps to the same replacement across every file in a corpus and the data stays useful
Run these algorithms over large volumes of data to prepare it for downstream agentic task building
Build adversarial test sets for de-identification across…