Data Engineer, Data Products
Data Engineer, Data Products
Location: New York City
Employment Type: Full time
Location Type: On-site
Department: Engineering
About Minerva
Minerva builds AI for marketing leaders. Our platform lets marketers focus on telling the story of their brand while AI agents handle the operationally intensive work: data management, analytics, campaign generation, measurement and reporting.
Everything is built on Minerva's proprietary consumer graph: an identity and attribute layer covering 270M+ U.S. consumers across 1,000+ through-time attributes. On top of it sit two agentic systems built in partnership with OpenAI: an Agentic Data Engineer that unifies and standardizes a brand's first-party data in hours, and an Agentic Data Scientist that trains robust targeting models at scale. Together, our data and platform improve the quality of a brand's first-party data, lift campaign performance and give marketing teams their time back. Our data team built Minerva's initial data product in less than a year and has already become best-in-class within the consumer data ecosystem.
We work with leading consumer brands across categories, including the NBA, Capital One, Hard Rock Stadium Group / Miami Dolphins, Wander and Trust & Will. We've raised $20M from The General Partnership, 8VC, Lingotto, NBA Investments, Topology Ventures, Future Positive, Background Capital and many others. Our team brings together operators and investors from Citadel, Dentsu, Bridgewater, Meta Superintelligence and Lazard, alongside researchers from Berkeley, MIT, Stanford and Cambridge.
About the Role
Minerva's data is not just infrastructure beneath our product; it is also one of our products. We are looking for a Data Engineer who can take ownership of complex consumer-data domains, develop a deep understanding of how their datasets relate and turn messy raw signals into trusted attributes and production data products. This role & Minerva are quite unique in the sense that GTM can immediately start selling your work output and generate enterprise-grade revenue in a matter of weeks.
You will spend most of your time at the transformation and derivation layer. You might become Minerva's internal expert on an identity graph, property and professional data, or a new source of consumer intent: learning the domain deeply enough to identify what is useful, what is misleading and what we should build next. Your job is not simply to make data available for someone downstream. You will use it to solve ambiguous problems and move an important part of our data product forward.
This is an end-to-end role. You will investigate novel datasets, source data when the answer is not already available, design domain models and derived attributes, and productionize your work so it can be consumed reliably by Minerva's applications, agents, APIs and ML models. Our data platform engineers build the infra that makes this work scalable; you must be comfortable operating within that infra and building your own ingestion and transformation pipelines without creating a bottleneck for the platform team.
The best fit can come from several backgrounds: a product-oriented data engineer, an applied data scientist who has data engineering skills, or a software engineer who has spent years solving difficult data problems. The common thread is first-principles reasoning, strong engineering fundamentals and a desire to own the answer from raw data through production. We test rigorously for data problem solving skills in our interview process.
What You'll Do
- Own one or more complex consumer-data domains end-to-end, becoming the person responsible for both understanding the data and advancing the products built from it.
- Investigate large, messy and unfamiliar datasets. Establish their grain, keys, relationships, coverage, failure modes and fitness for different product use cases.
- Design and build durable domain models, derived attributes and entity relationships that can power Minerva's applications, AI agents, APIs, customer deliveries and predictive models.
- Build and operate the ingestion and transformation pipelines required to bring your work into production, including validation, observability, backfills and recovery. Productionization is key.
- Go on data quests: identify and evaluate new sources, determine how they can improve our consumer graph and find clever ways to extract signal from imperfect inputs. We often have budget for purchasing new data when there is clear ROI.
- Make data outputs trustworthy enough to be consumed autonomously. Define quality checks, provenance and guardrails that distinguish reliable signal from convenient but misleading data.
- Partner with data scientists, platform engineers, product engineers and customer-facing teams to turn open-ended business or product questions into scalable data products.
- Use LLMs, embeddings and modern AI development tools where they materially improve data standardization, classification, enrichment or engineering velocity.
- Find new ways to create and protect business value through Minerva's proprietary data asset, from improving existing products to opening entirely new revenue opportunities.
Our Data Stack
- Dagster for orchestration
- dbt-core within Dagster as a primary data-transformation surface
- Snowflake for analytical workloads
- Spark, Iceberg, Trino and AWS Glue for lakehouse workloads
- Postgres, Elasticsearch and other product-facing systems
- Frontier and open-source models, agent SDKs and batch APIs from OpenAI and Anthropic
Qualifications
- 2-5+ years working as a data engineer, software engineer or applied data scientist in a data-heavy context. Your prior title matters less than evidence that you live and breathe data.
- Highly proficient in Python and SQL.
- Driven by first-principles thinking. You can take an ambiguous data problem, determine what must be true, interrogate the available evidence and design a practical path to an answer.
- Strong intuition for data cleaning, ingestion and data modeling. We expect these foundations to be second nature so your thinking is free for larger and more ambiguous data initiatives, especially given the leverage of modern AI coding tools.
- Comfortable building and deploying production data pipelines, not just analyzing data in notebooks or handing specifications to another engineering team.
- Able to balance analytical depth with engineering pragmatism: you care whether an attribute is conceptually valid and whether it can be produced reliably at scale; you know when a new idea won't provide any lift.
- Comfortable owning an ambiguous initiative end-to-end in a lean, fast-changing environment. You're more product-minded than people give you credit for.
- Willingness to work in our New York City office. We provide a relocation package.
- Eagerness to learn, grow and raise the bar with your coworkers.
Preferred
- Experience working with large, messy, multi-source datasets where the semantics were not obvious and documentation was incomplete.
- Experience with consumer data, identity resolution, entity graphs, property data, behavioral or intent data, or other complex third-party data domains.
- Experience with orchestration tools such as Dagster, Airflow or Prefect and transformation tools such as dbt or SQLMesh.
- Experience with analytical databases such as Snowflake, Redshift or BigQuery and familiarity with transactional databases such as Postgres or MySQL.
- Experience with lakehouse or distributed-processing systems such as Spark, Iceberg, Trino or AWS Glue.
- Familiarity with AWS or another major cloud platform.
- Exposure to ML eng/ops, applied ML or feature engineering. Deep modeling expertise is not required.
- Effective use of AI coding tools such as Claude Code, Cursor or OpenCode as a force multiplier.
- Prior experience at an early-stage startup.
You do not need to tick every box. If you are a strong with data, we want to hear from you.
Compensation
Base salary: $200,000 to $225,000, commensurate with experience. Competitive equity and a marquee benefits package.