- Design, build, and operate scalable Extract, Transform, Load / Extract, Load, Transform (ETL/ELT) data pipelines (batch and real-time) to ingest, transform, and load data from multiple sources.
- Engineer and maintain data platforms (data lakes and data warehouses) with strong reliability, availability, and operational resilience.
- Implement system integrations using Application Programming Interfaces (APIs), streaming, and enterprise data integration tools to ensure consistent end-to-end data flows.
- Apply Apache Spark / Python for Spark (Spark/PySpark) and advanced Structured Query Language (SQL) for large-scale processing, optimization, and perform data transformations.
- Embed data quality, governance, security, and regulatory controls across pipelines and datasets.
- Enable development and operations (DevOps) for data through Continuous Integration / Continuous Delivery (CI/CD), orchestration and scheduling, monitoring, troubleshooting, and production support.
- Hands-on expertise in Prophecy pipeline development and deployment, PySpark analysis, and PySpark pipeline engineering (build and release), alongside strong core software engineering fundamentals
Key Skills : Python, Spark, Apache, ETL, data pipelines, API's.
Graduate in Computer Science, Data Science, or related field. 2-3 years of experience in data engineering or related field.