Data Engineer - Consultant
Rubick · Jayanagar, Bengaluru / Remote · India · Remote
Posted Jul 29, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Data Engineer - Consultant
Role Overview
You will be responsible for building and maintaining Rubick's data acquisition engine by extracting, validating, and structuring product information from eCommerce websites and marketplaces. This role focuses on developing scalable web crawling solutions that power Product Discovery, Search, Catalog Intelligence, and Market Intelligence platforms.
The role involves three areas: Part 1: The Fundamentals | Part 2: AI-Driven Data Crawling Excellence | Part 3: Innovation & Improvements
Part 1: The Fundamentals
Develop, maintain, and optimize web crawlers to extract product data from eCommerce websites and marketplaces.
Collect structured and unstructured product information, including product details, pricing, images, specifications, and availability.
Clean, validate, and organize extracted datasets to ensure accuracy and consistency.
Monitor crawler performance, identify failures, and resolve data extraction issues.
Collaborate with Product, Engineering, and Data teams to ensure reliable and timely data collection.
Maintain documentation for crawling processes, extraction rules, and data quality standards.
Part 2: AI-Driven Data Crawling Excellence
Leverage AI-powered extraction techniques to improve data accuracy and extraction efficiency.
Build intelligent crawling workflows using Python and modern web automation frameworks.
Develop scalable solutions for handling dynamic websites, JavaScript-rendered content, and anti-bot mechanisms.
Utilize browser automation, APIs, proxies, and scheduling tools to maximize crawl success rates.
Implement automated data validation and monitoring systems to maintain high-quality datasets.
Collaborate with AI and Data Engineering teams to support Machine Learning models and Product Intelligence systems.
Part 3: Innovation & Improvements
Continuously optimize crawling performance for speed, scalability, and reliability.
Identify opportunities to automate repetitive extraction and validation workflows.
Improve data collection strategies by adopting new crawling technologies and AI-assisted solutions.
Build reusable crawling frameworks and standardized extraction pipelines.
Stay updated with the latest web scraping libraries, browser automation tools, and industry best practices.
Contribute to knowledge repositories, documentation, and process improvements across the data acquisition function.
To Have
1+ year of experience in Web Crawling, Data Scraping, Product Matching, Data Extraction, or similar roles.
Strong proficiency in Python for web scraping and automation.
Hands-on experience with BeautifulSoup, Scrapy, Selenium, Playwright, or similar frameworks.
Good understanding of HTML, CSS, XPath, JSON, and DOM structures.
Familiarity with REST APIs, proxies, browser automation, and dynamic website scraping.
Strong analytical and problem-solving skills.
Ability to work with large datasets while maintaining…