Overview
The Data Engineer role focuses on designing and implementing a robust data architecture using Azure and Databricks within the financial services sector. The contractor will collaborate with team members to develop data ingestion pipelines and establish governance mechanisms, enabling efficient processing and reporting of data. This position is fully remote and aims to optimize data workflows while ensuring high standards of quality and compliance.
Responsibilities
- Define and build a data lake architecture on Azure utilizing ADLS and Databricks.
- Develop ingestion patterns from API sources, both batch and streaming.
- Establish data governance practices, including metadata management and access controls.
- Define storage structures, naming conventions, and schema design for data organization.
- Develop API connectors and ingestion pipelines using Python.
- Implement data transformation logic using PySpark/Python for cleansing, standardization, and validation.
- Configure monitoring, logging, and alerting mechanisms for data pipelines.
- Execute unit and integration testing for data workflows.
Requirements
- Proven experience as a Data Engineer with strong knowledge of Azure and Databricks.
- Proficiency in Python for developing data pipelines and connectors.
- Experience with data lake architectures and layered data models.
- Familiarity with API integration and data ingestion methods.
- Knowledge of data governance, metadata management, and access control systems.
- Experience with data transformation tools, particularly PySpark.
- Exposure to Financial Services or Banking sector is a plus but not required.