Overview
The Senior Data Engineer will play a pivotal role in shaping our client's data strategy while working closely with the Head of Data and Lead Data Scientist. The contractor will focus on translating business objectives into high-value data projects and is expected to execute these strategies through production-grade coding and pipeline development.
Responsibilities
- Partner with the Head of Data and Lead Data Scientist to define the data lakehouse strategy and schema evolution.
- Translate Objectives and Key Results (OKRs) into detailed Technical Design Documents (TDDs) that meet compliance standards.
- Establish data governance, lineage tracking, and multi-tenant isolation for workflows.
- Define JIRA Epics and user stories based on approved TDDs for implementation.
- Architect, code, and deploy idempotent, backfill-safe data pipelines using Python and Dagster.
- Utilize AWS cloud environments to enforce data lineage and schema evolution across storage layers.
Requirements
- 4+ years of experience in Python distributed data engineering and systems backend development.
- Hands-on production experience with Airflow in a Medallion or data lakehouse environment.
- Proficiency in Terraform for managing AWS resources like EKS and Aurora PostgreSQL.
- Familiarity with enterprise audit standards such as SOC2 or ISO27001.
- Experience with multi-agent frameworks or libraries like Google ADK or Langchain is a plus.
- Knowledge of RAG pipelines, pgvector embeddings, or LLM inference pipeline deployment is advantageous.