JobRaahGet matched free

Jobs

Data Engineer

Riseup Labs

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

Job Context: We are seeking a senior Data Engineer to design, build, and operate scalable data pipelines and curated data assets that enable AI-driven workflows and business-critical automation. This role is focused on production-grade delivery, data quality, governance, and integration across enterprise systems, supporting outcomes such as cycle time reduction, productivity gains, cost reduction, and improved operational decisioning. The Data Engineer will work closely with architecture, platform engineering, AI engineers, and business stakeholders to ensure data is accessible, reliable, secure, and fit for operational and AI-enabled use cases. Job Responsibilities: Design and implement scalable data pipelines (batch and near real-time) to support AI and automation use cases. Build and maintain robust ingestion processes across structured, semi-structured, and unstructured sources. Implement data transformation, cleansing, validation, and quality controls aligned with enterprise standards. Provide curated, reusable datasets and interfaces to enable AI use cases, including ingestion and indexing flows used for retrieval-based AI patterns (RAG) where applicable. Define and enforce data contracts, schema management, and versioning to support reliable downstream consumption. Implement data preparation logic for RAG, including semantic chunking strategies, metadata enrichment, and approaches for handling incremental updates and re- indexing. Collaborate with AI Engineers to define indexing schemas and metadata strategies that optimize retrieval performance. Collaborate with AI Engineers to ensure source data and retrieval pipelines are production-ready, measurable, and aligned with delivery objectives. Implement observability for pipelines (freshness, completeness, accuracy, lineage signals where applicable). Troubleshoot pipeline failures and performance bottlenecks and ensure stable production operations. Ensure alignment with internal governance requirements (access control, data privacy constraints, auditability). Educational Requirements: Bachelor’s or Master’s in Computer Science, Software Engineering, Information Systems, Data Engineering, or related field. Experience with data pipelines, SQL, cloud platforms is key. Required Skills and Experience: Strong experience building and maintaining data pipelines in production environments. Expert-level SQL skills and solid understanding of enterprise data modeling concepts. Advanced Python experience for pipeline development and automation. Experience processing semi-structured and unstructured data (for example PDFs, JSON payloads, Markdown/text corpora), including preparing data for consumption by LLM- enabled applications and retrieval systems. Experience working with APIs, enterprise systems, and heterogeneous data sources. Strong understanding of data quality patterns (validation rules, anomaly detection, automated checks). Ability to operate with high ownership and…