Data Engineer – Applied ML
Similarweb · Tel Aviv, Israel · On-site
Posted Sep 30, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Similarweb is the leading digital intelligence platform used by over 3500 global customers. Our wide range of solutions power the digital strategies of companies like Google, eBay, and Adidas.
We help our customers succeed in today’s digital world by giving them access to data-driven insights, competitive benchmarks, strategic analysis, and more.
In 2021, we went public on the New York Stock Exchange, and we haven’t stopped growing since!
We’re looking for a Data Engineer with a strong applied ML focus to join our R&D department!
Why is this role so important at Similarweb?
Our Retail Intelligence products help leading brands and retailers understand how their products, brands and categories perform online. Behind them is data collected from retailers and marketplaces around the world: product pages, brands and categories, each described differently by every site.
Your mission is to turn that data into a single, trusted view: classifying products into a unified taxonomy, normalizing brands and attributes, and matching the same entities across sources. And doing it at scale, across a catalog of more than a billion records that keeps growing and changing every day.
This is an applied ML role within data engineering. You’ll build with LLMs, agentic frameworks such as LangGraph, embeddings and classical ML, and ship them as production pipelines. It’s hands-on work, not research for its own sake, but it takes a real understanding of classification and NLP methods to choose the right tool for each problem and prove that it works.
So, what will you be doing all day?
Your daily responsibilities may include:
Designing and building LLM-powered and ML-based pipelines that classify, normalize, structure and match product, brand and category data
Building agentic workflows (LangGraph or similar) that automate complex data tasks end to end
Choosing the right approach for each problem (LLMs, embeddings, fine-tuned models, classical classifiers or rules), balancing accuracy, cost and latency
Scaling solutions to run efficiently over billions of records, using Spark, Databricks and our cloud infrastructure
Building evaluation frameworks: ground-truth datasets, labeling processes, quality metrics and ongoing monitoring
Taking solutions from POC to production, and owning them after launch
Working closely with Product to define requirements and shape the roadmap
Collaborating with data engineers, data scientists and other R&D teams on infrastructure and best practices
This is the perfect job for someone who:
Holds a B.Sc. or M.Sc. in Computer Science, Data Science, Mathematics or another relevant field
Has 4+ years of hands-on experience as a data engineer, ML engineer or data scientist, with solutions running in production
Has strong Python skills and writes production-quality code
Has hands-on experience building LLM-based applications in production (prompt engineering, structured outputs, RAG, embeddings,…