Design & Build petabyte-scale data pipelines — Design, develop, and operate reliable ingestion and transformation pipelines. Proven ability to design system architecture for large-scale data platforms — data modeling, layering, scalability, reliability, and cost trade-offs. Convert raw data into intelligence — Transform staging → canonical → curated data products, building purpose-built, curated semantic/logical models. Mine logs to drive optimization — Instrument, collect, and analyze telemetry and operational logs to surface anomalies, root-cause trends, and process-optimization opportunities. Design & build agents for leadership and other personas — Develop AI-native agents that decompose natural-language questions, route them to the right domain skill, and return insights driven by reasoning, insights, and recommendations. Deliver KPIs that measure business outcomes — Partner with stakeholders to define, model, and ship the KPIs and persona dashboards that quantify business impact and inform executive decision-making. Rapidly prototype and prioritize ideas Strong data modeling skills — dimensional/semantic modeling, schema design, normalization/denormalization trade-offs, and building reusable curated models across domains. Own engineering fundamentals — Enterprise-grade logging, telemetry, evaluations, regression, performance testing, security, and continuous learning across every layer of the platform. Advance data governance — Build with data tagging, RBAC/attribute-grained policy, lineage, audit telemetry, compliance, and data-quality standards baked in. Design-first: you think in architecture — clear contracts, reusable models, and patterns that scale — and communicate them through crisp design docs and reviews. Required 3-5 years of software/data engineering experience building production data pipelines at scale. Strong programming skills (e.g., Python, and/or Scala) and expert SQL. Hands-on experience with large-scale distributed data processing (Spark/Databricks, Synapse) and cloud data lakes (ADLS Gen2 / Fabric OneLake). Experience designing curated/semantic data models and delivering KPIs to business stakeholders. Solid grounding in data governance, security, lineage, and data quality. Ability to translate ambiguous business questions into robust, well-modeled data products. Experience building LLM/agentic applications — orchestration, RAG/grounding, NL2SQL/NL2DAX, skill routing, or MCP-based integrations. Familiarity with Azure AI Foundry, Copilot Studio, and evaluation/fine-tuning workflows for agents.
Want jobs like this matched to you?
SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.