Lead Data Engineer - Enterprise Data & Analytics - Remote

Rochester, MNFull-timePosted Jul 31, 2026

Lead data design, prototype, and development of data pipeline architecture pipelines. Lead implementation of internal process improvements: automating manual processes, optimizing data delivery, re-designing infrastructure for greater scalability. Lead cause analysis on external and internal processes and data to identify opportunities for improvement and answer questions. Excellent analytic skills associated with working on unstructured datasets. Understand the architecture, be a team player, lead technical discussions and communicate the technical discussion. Be a senior Individual contributor of the Data or Software Engineering teams. Be part of Technical Review Board along with Manager and Principal Engineer. Be a technical liaison between Manager, Software Engineers and Principal Engineers. Collaborate with software engineers to analyze, develop and test functional requirements. Serve as a hands-on technical leader who actively designs, develops, reviews, and optimizes production-grade data pipelines, data products, and platform capabilities. Maintain significant contribution to production codebases while establishing engineering standards, mentoring team members, and driving delivery of scalable, resilient solutions. Mentor and Coach Engineers. Work with team members to investigate design approaches, prototype new technology and evaluate technical feasibility. Work in an Agile/Safe/Scrum environment to deliver high quality software. Establish architectural principles, select design patterns, and then mentor team members on their appropriate application. Facilitate and drive communication between front-end, back-end, data and platform engineers. Play a formal Engineering lead role in the area of expertise. Keep up-to-date with industry trends and developments.


Key Responsibilities:
These positions are hands-on engineering roles. In this role, employees are expected to actively design, develop, review, and optimize production code and platform capabilities while providing technical leadership and mentorship to engineering teams.

Bachelor’s Degree in Computer Science/Engineering or related field with 6 years of experience OR an Associate’s degree in Computer Science/Engineering or related field with 8 years of experience. Knowledge of professional software engineering practices and best practices for the full software development life cycle (SDLC), including coding standards, code reviews, source control management, build processes, testing, and operations. Have in-depth knowledge of data engineering and building data pipelines with a minimum of 5 years of experience in data engineering, data science or analytical modeling and basic knowledge of related disciplines. Worked and lead Data Engineering teams in Continuous Integration / Continuous Delivery model. Build/Lead Data products highly resilient in nature. Build/Lead Test Automation suites, Unit Testing coverage, Data Quality, Monitoring & Observability. A minimum experience of 5 years using relational databases and NoSQL Databases. Experience with cloud platforms such as GCP, Azure, AWS.
Continuous Integration using Jenkins, Git Hub Actions or Azure Pipelines. Experience with cloud technologies, development and deployment. Experience with tools like Jira, GitHub, SharePoint, Azure Boards. Experience using advanced data processing solutions/capabilities such as Apache Spark, Hive, Airflow and Kafka, GCP Dataflow. Experience using big data, statistics and knowledge of data related aspects of machine learning. Experience with Google BigQuery, FHIR APIs, and Vertex AI. Knowledge of how workflow scheduling solutions such as Apache Airflow and Google Composer related to data systems. Knowledge of using Infrastructure as code (Kubernetes, Docker) in a cloud environment.

 

The preferred candidate will possess:

  • Advanced proficiency in Python and SQL with demonstrated experience building and supporting production-grade solutions.
  • Advanced experience designing and implementing scalable distributed computing solutions using technologies such as Spark, Flink, Ray, or comparable frameworks.
  • Deep understanding of cloud-agnostic architecture principles and modern data platform design.
  • Advanced experience with open data architecture technologies including Apache Iceberg, Delta Lake, and Apache Hudi.
  • Strong expertise with modern analytical data formats including Parquet, Avro, and ORC.
  • Experience designing data platforms that support analytics, AI/ML, and operational workloads at enterprise scale.
  • Experience implementing CI/CD, automated testing, Infrastructure-as-Code, observability, and engineering best practices.
  • Experience designing systems for scalability, reliability, security, resiliency, and long-term maintainability.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free