Senior Staff Data Engineer, Databricks (R5492)
Job Description:
The Senior Staff Data Engineer will help build and operate the enterprise lakehouse on Databricks, creating the governed data foundation that supports multiple business domains and downstream analytics. This is a hands-on role, responsible for scalable ingestion, reliable data processing, and strong technical controls across the Bronze and Silver layers of the medallion architecture.
This role is not limited to moving data from point A to point B. The Data Engineer is expected to understand the meaning, sensitivity, classification, and intended use of the data flowing through the pipelines they build, and to apply that understanding when designing controls, access patterns, data quality checks, and segregation boundaries appropriate for a highly regulated environment.
Job Description:
The Senior Staff Data Engineer will help build and operate the enterprise lakehouse on Databricks, creating the governed data foundation that supports multiple business domains and downstream analytics. This is a hands-on role, responsible for scalable ingestion, reliable data processing, and strong technical controls across the Bronze and Silver layers of the medallion architecture.
This role is not limited to moving data from point A to point B. The Data Engineer is expected to understand the meaning, sensitivity, classification, and intended use of the data flowing through the pipelines they build, and to apply that understanding when designing controls, access patterns, data quality checks, and segregation boundaries appropriate for a highly regulated environment.
What you'll do:
- Design and build ingestion pipelines, batch and streaming where appropriate, from enterprise source systems into the Databricks lakehouse using Delta Lake.
- Own Bronze-layer ingestion, including raw landing patterns, metadata capture, load traceability, and recoverable ingestion design.
- Build Silver-layer pipelines for cleansing, standardization, deduplication, conformance, and quality enforcement without embedding unauthorized or undocumented KPI logic.
- Define and evolve reusable ingestion and transformation patterns, templates, and engineering standards that other domains can adopt as they onboard to the platform.
- Implement and maintain Databricks platform constructs needed for secure delivery, including catalogs, schemas, service principals, job orchestration, and environment-aware deployment patterns.
- Build and maintain CI/CD pipelines for data platform assets, including ingestion code, transformation logic, workflow definitions, tests, and environment promotion across dev, test, and prod.
- Apply data classification, segregation, and handling requirements within the pipeline design, ensuring sensitive and regulated data is processed in accordance with enterprise controls and access policies.
- Build data quality controls that do more than detect technical failures, including checks that reflect actual business meaning, record integrity, completeness, and expected domain behavior.
- Quarantine, flag, and route problematic records or datasets according to defined quality and compliance rules rather than silently dropping or obscuring issues.
- Maintain documentation for source objects, ingestion logic, applied transformations, data quality rules, and known limitations so downstream teams can trust and use the data correctly.
- Partner with the Analytics Engineer and domain teams to ensure Silver-layer data is reliable, well-governed, and suitable for trusted Gold-layer modeling.
- Collaborate with domain engineering teams to align on ownership boundaries, onboarding patterns, data contracts, and support expectations as new domains are enabled onto the platform.
Required qualifications:
- 12+ years of data engineering experience, including hands-on ownership of production data pipelines.
- Strong Databricks experience, including Delta Lake, Databricks Workflows or Jobs, and Spark with PySpark and/or Spark SQL.
- Working knowledge of Unity Catalog, including catalogs, schemas, tables, lineage, and access control concepts.
- Experience with batch, CDC, and/or streaming ingestion patterns and the operational trade-offs associated with each.
- Experience with CI/CD and deployment automation for data pipelines and platform assets, including version control, testing, and controlled promotion across environments.
- Strong SQL skills and solid grounding in data modeling fundamentals, even if dimensional modeling is not the primary responsibility of this role.
- Demonstrated ability to understand the business and regulatory context of the data being processed, not just the mechanics of pipeline development.
- Experience applying data classification, segregation, security, retention, or compliance requirements in data engineering workflows within a regulated or security-sensitive environment.
- Ability to design pipelines with awareness of the actual data domains involved, including sensitivity, ownership, permitted use, and downstream impact.
- Comfort operating in a fast-moving platform build where patterns are still being established and engineers are expected to shape standards, not just follow them.
Preferred qualifications:
- Databricks certification such as Data Engineer Associate or Professional.
- Experience standing up or maturing Unity Catalog structures and access models in an enterprise setting.
- Experience with CI/CD, infrastructure-as-code, and deployment automation for data platforms.
- Experience in defense, aerospace, financial services, healthcare, or another highly regulated environment.
- Experience enabling multiple business domains on a shared data platform while maintaining strong governance and ownership boundaries.
Full-time regular employee offer package: Pay within range listed + Bonus + Benefits + Equity Temporary employee offer package: Pay within range listed above + temporary benefits package (applicable after 60 days of employment) Salary compensation is influenced by a wide array of factors including but not limited to skill set, level of experience, licenses and certifications, and specific work location. All offers are contingent on a cleared background and possible reference check. Military fellows and part-time employees are not eligible for benefits. Please speak to your talent acquisition representative for more information. ### Shield AI is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, marital status, disability, gender identity or Veteran status. If you have a disability or special need that requires accommodation, please let us know.