Senior Data Engineer - Gurugram
CLANX · Gurgaon (Gurugram), Haryāna, India · Hybrid
Posted Jul 3, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Mechademy is hiring a Senior Data Engineer to build and scale reliable data platforms, pipelines, and models that power enterprise AI, machine learning, and analytics for industrial asset monitoring and predictive maintenance.
Company Details
Mechademy is an enterprise AI company building real-time monitoring, diagnostics, and predictive maintenance solutions for industrial equipment. The company serves clients across oil & gas, power generation, and LNG sectors through production-grade AI and physics-informed machine learning systems.
Website: https://mechademy.com/
Responsibilities
What You’ll Own
Lakehouse Pipelines & Ingestion (35%)
Design and own batch ETL/ELT and CDC pipelines that bring sensor and operational data into the lakehouse, orchestrated in Dagster
Build for reliability: idempotent, incremental, backfill-safe pipelines with sane retry and failure handling, that still produce correct output when a worker is killed mid-run or a message is delivered twice
Onboard new client data sources: schema and tag mapping, time-series normalization, resampling, gap handling at scale
2. Modeling & Serving (25%)
Model raw data into well-structured, documented tables that downstream ML and analytics can trust
Build and maintain the datasets behind ML feature pipelines and the lakehouse layer powering self-serve analytics
Write performant Spark/PySpark and SQL; optimize partitioning, storage formats, and query cost
3. Data Quality & Reliability (10%)
Own data quality: validation, freshness/SLA monitoring, and observability so bad data is caught before it reaches consumers
Make the data layer debuggable: lineage, tests, and alerting that tell you what broke and where
Reason about failure modes across the whole path (queue, worker, orchestrator, database, object store) and design so that a partial failure leaves the system in a state you can recover from
4. Relational & Operational Data (30%)
Contribute to the schema, indexing, and query performance of the relational database the product runs on
Design tables and constraints so that correctness is enforced at the database layer, and diagnose slow queries from their plans
Own retention and the boundary between the operational database and the lakehouse: what stays, what moves, and how it gets there
What Success Looks Like
First 30 days: Productive in the codebase and orchestration layer. First pipeline change merged.
First 90 days: Independently shipping and owning pipelines. Onboarded at least one new data source end-to-end.
First 6 months: Owning a lakehouse data domain, its ingestion, models, and quality, that ML and analytics teams rely on you to drive.
Requirement
Must-Have
4+ years building production data pipelines: real systems with real consumers, not just one-off scripts
Strong data engineering fundamentals: data modeling, batch vs. streaming, idempotency, incremental processing, partitioning.
Expert SQL and strong Python: query optimization,…