Data Engineer

Overview

The Data Engineer will focus on creating a data lake architecture on Azure using Databricks, emphasizing data ingestion, processing, and reporting. This fully remote position for a 3-month contract will involve collaboration with various teams to establish data governance and develop efficient API connectors. The role is essential for ensuring that data flows smoothly and adheres to defined standards and governance.

Responsibilities

  • Define and build a data lake architecture on Azure using ADLS and Databricks.
  • Develop ingestion patterns from API sources for both batch and streaming data.
  • Establish data governance, metadata management, and access controls.
  • Define storage structures, naming conventions, and schema designs.
  • Develop API connectors and ingestion pipelines using Python.
  • Implement data transformation logic using PySpark and Python for data cleansing and validation.
  • Configure monitoring, logging, and alerting for data workflows.
  • Execute unit and integration testing across data pipelines.

Requirements

  • Proven experience with Azure Databricks.
  • Strong proficiency in Python programming.
  • Experience in building data ingestion pipelines.
  • Familiarity with data governance frameworks and metadata management.
  • Understanding of data lake architecture and layering (bronze, silver, gold).
  • Experience with data transformation using PySpark is a plus.
  • Desirable: Exposure to Financial Services, Insurance, or Banking.