Overview
We are seeking an experienced Data Engineer to join a cutting-edge technology program focused on building a scalable data platform for advanced AI applications and maritime surveillance. In this remote role, you will work within an autonomous engineering team to enhance a proven prototype into a production-grade platform, processing large volumes of complex real-world data. A strong background in Python and expertise in building distributed data pipelines are essential for success in this position.
Responsibilities
- Design and productionise scalable, high-throughput distributed data pipelines.
- Onboard and normalise diverse data sources into consistent, queryable schemas.
- Build the core data foundations used by AI/ML models and agents.
- Develop a knowledge layer allowing AI agents to read and write data with full traceability back to source.
- Work with complex IoT, geospatial, time-series, and maritime data.
- Build secure, high-performance production systems using Python.
- Help shape technical and architectural decisions in a greenfield environment.
Requirements
- Strong production-level experience with Python.
- Demonstrated ability to build distributed data pipelines at scale.
- Solid understanding of data modeling, schema design, and data lineage/provenance.
- Strong software engineering fundamentals.
- Experience in taking data systems from prototype to production.
- Eligibility or willingness to obtain SC clearance.
- Desirable experience with Kafka, Flink, or other streaming technologies.
- Familiarity with Docker and Kubernetes.