JobRaahGet matched free

Jobs

Big Data Engineer

Wingify · Remote / Delhi, India · Remote

Posted Aug 18, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About the Role We are looking for a Big Data Engineer to design, build, and operate large-scale data pipelines and analytical infrastructure that transform high-volume raw data into reliable, query-ready datasets for analytics, reporting, and data-driven products. Our data platform ingests and processes data from multiple sources and serves analytics, data science, product, and downstream applications. In this role, you will own data pipelines end-to-end—from ingestion and transformation to warehousing, orchestration, data quality, and observability. A key part of the role is owning ClickHouse as our primary analytical data store . You will be responsible for designing scalable data models, optimizing query performance, and ensuring the platform remains reliable and cost-efficient as data volumes and workloads grow. You will work closely with data scientists, analysts, product engineers, and other engineering teams to build a modern, cloud-native data platform. What You'll Do Design, build, and maintain robust batch and streaming data pipelines that ingest data from multiple sources into analytical data stores. Build and operate Apache Airflow DAGs , including scheduling, dependencies, retries, backfills, idempotency, concurrency, and failure handling. Develop analytics-ready datasets using dbt , following well-structured staging, intermediate, and mart layers with appropriate tests, documentation, and incremental models. Own ClickHouse as the primary analytical store, including: Schema and table design using the MergeTree family of engines Partitioning and sorting/primary key strategies Materialized views Distributed and replicated table architectures Query and memory optimization High-volume data ingestion and performance tuning Work with BigQuery where cloud data-warehouse patterns are appropriate, including data modeling and query/cost optimization. Design and operate NoSQL and key-value data stores , including Bigtable, DynamoDB, and Redis, based on specific access patterns and performance requirements. Build and maintain data-quality frameworks covering validation, testing, freshness, completeness, reconciliation, and anomaly detection. Implement monitoring, alerting, structured logging, and observability for data pipelines and services. Own pipeline SLAs, incident response, troubleshooting, and root-cause analysis. Manage backfills, safe re-runs, schema evolution, and data migrations while minimizing downstream impact. Build reproducible, containerized environments using Docker and contribute to CI/CD and Infrastructure as Code practices. Partner with analysts, data scientists, product managers, and product engineers to translate business and technical requirements into scalable data models and pipelines. Continuously improve pipeline reliability, scalability, performance, and infrastructure cost efficiency. Must-Have Requirements 4-6 years of experience in data engineering or a closely related field, with…