Principal Data Engineer - Enterprise Data & Analytics - Remote
The Principal Data Engineer serves as a hands-on technical authority responsible for defining and implementing enterprise-scale data architecture and engineering strategies while actively contributing to solution design, development, optimization, and technical delivery. As part of an assigned product team, this role develops and deploys data pipelines, integrations, and transformations to support analytics and machine learning applications using open-source programming languages and vendor software. The position requires a strong understanding of the organization’s current solutions, coding languages, tools, and Enterprise Data and Analytics technology framework, as well as the ability to apply independent judgment, provide consultative services to departments, divisions, and leadership committees, and partner with product owners and Analytics and Machine Learning delivery teams to identify and retrieve data, conduct exploratory analysis, transform data, visualize trends, build and validate analytical models, and translate qualitative and quantitative assessments into actionable insights.
Key responsibilities:
These positions are hands-on engineering roles. In this role, employees are expected to actively design, develop, review, and optimize production code and platform capabilities while providing technical leadership and mentorship to engineering teams.
A Bachelor's degree in a relevant field such as engineering, mathematics, computer science, information technology, health science, or other analytical/quantitative field and a minimum of seven years of professional or research experience in data visualization, data engineering, analytical modeling techniques; OR an Associate's degree in a relevant field such as engineering, mathematics, computer science, information technology, health science, or other analytical/quantitative field and a minimum of nine years of professional or research experience in data visualization, data engineering, analytical modeling techniques. In-depth business or practice knowledge will also be considered.
Incumbent must have the ability to manage a varied workload of projects with multiple priorities and stay current on healthcare trends and enterprise changes. Interpersonal skills, time management skills, and demonstrated experience working on cross functional teams are required. Requires strong analytical skills and the ability to identify and recommend solutions and a commitment to customer service. The position requires excellent verbal and written communication skills, attention to detail, and a high capacity for learning and problem resolution. Advanced experience in SQL is required. Advanced Experience in scripting languages such as Python, JavaScript, PHP, C++ or Java & API integration is required. Experience in hybrid data processing methods (batch and streaming) such as Apache Spark, Hive, Pig, Kafka is required. Experience with big data, statistics, and machine learning is required. The ability to navigate linux and windows operating systems is required. Knowledge of workflow scheduling (Apache Airflow Google Composer), Infrastructure as code (Kubernetes, Docker) CI/CD (Jenkins, Github Actions) is required. Experience in DataOps/DevOps and agile methodologies is required. Experience with hybrid data virtualization such as Denodo is preferred. Working knowledge of Tableau, Power BI, SAS, ThoughtSpot, DASH, d3, React, Snowflake, SSIS, and Google Big Query is preferred.
The preferred candidate will possess:
- Expert-level proficiency in Python and SQL with extensive experience developing enterprise-scale production systems.
- Advanced expertise in scalable distributed computing frameworks and modern data processing platforms.
- Advanced experience implementing and governing open data architectures utilizing Apache Iceberg, Delta Lake, Apache Hudi, and related technologies.
- Deep understanding of modern analytical storage formats including Parquet, Avro, and ORC.
- Demonstrated expertise in lakehouse architecture, data platform design, and large-scale data engineering practices.
- Experience architecting and implementing cloud-agnostic solutions across multiple technology ecosystems.
- Experience designing highly scalable, fault-tolerant, secure, and observable data platforms supporting analytics, AI, machine learning, and operational workloads.
- Experience establishing enterprise engineering standards, architecture patterns, and modernization strategies.