Data Engineer III (Batch Processing)
At Expedia Group, we help travelers explore the world, one journey at a time. As a global travel company powered by passionate people, trusted partnerships, and leading technology, we connect travelers, partners, and advertisers through our consumer brands, B2B network, and travel advertising business.
Here, you'll do meaningful work that helps millions of people discover, book, and experience travel with more ease, confidence, and joy. Our five Behaviors-Traveler First, Think Big, Operate with Excellence, Ownership Mindset, and Succeed Together-help foster a supportive environment where people can grow their careers and have the flexibility, benefits, and support to do their best work. Join us and build for travelers everywhere.
Data Engineer III (Batch Processing)
Introduction to the team
Our Technology team partners across Expedia Group to create products, services, and tools that deliver high-quality experiences for travelers, partners, and employees.
Within the Data Platform organization, the EG Metrics Platform (EGMP) team is building Expedia Group’s enterprise-wide semantic layer: a centralized metrics platform that standardizes how business metrics are defined, computed, governed, and consumed across the company. As the single source of truth for critical business metrics such as bookings, revenue, visits, funnel progression, and conversion, EGMP powers executive reporting, experimentation, BI dashboards, and self-serve analytics at scale.
Semantic layers are increasingly important in the age of LLMs and AI agents because they improve consistency, reduce hallucinations, and enable trusted metric discovery. Our team is building toward an AI-powered semantic layer with agentic capabilities for natural-language metric exploration, intelligent anomaly detection, automated operations, and governed AI-ready data access.
About the role
We are looking for a Data Engineer III who combines strong data engineering fundamentals with solid hands-on experience in Apache Spark, workflow orchestration, and modern data platform development. This role is ideal for an engineer who can independently build and optimize production data pipelines, contribute to shared platform capabilities, and work effectively with cross-functional partners.
You will help design and operate data products that power metric computation, analytics, and AI-assisted workflows across Expedia Group. You should be comfortable working with batch data pipelines, orchestration frameworks, SQL-based transformations, and data access patterns that support downstream consumers.
In this role, you will
- Design, build, and operate scalable, reliable data pipelines across large-scale distributed data platforms.
- Develop and optimize Spark-based data processing jobs, with a strong understanding of performance tuning, partitioning, joins, shuffles, and troubleshooting production issues.
- Write and optimize advanced SQL across engines such as Spark SQL, Trino/Presto, and similar distributed query systems for large-scale transformations, aggregations, and analytics use cases.
- Build and maintain analytics-ready data models including fact tables, dimensions, aggregates, metric layers, and gold datasets used by dashboards, experimentation platforms, and executive reporting.
- Build and maintain Airflow workflows for dependable scheduling, monitoring, recovery, and operational support.
- Contribute to data access services, APIs, or integration interfaces that help downstream applications and tools consume trusted metrics and datasets.
- Contribute to platform capabilities for metric discovery, SQL generation, anomaly detection, and AI-assisted operations, including agent skills, MCP tools, and workflow automation.
- Use AI-assisted engineering tools in day-to-day development to improve productivity in coding, debugging, documentation, testing, and operational analysis.
- Develop data quality, observability, and alerting capabilities that improve operational reliability across production workflows.
- Partner with analysts, data scientists, product managers, and engineers to translate business needs into scalable technical solutions and self-serve data capabilities.
- Own medium-to-large deliverables end to end, from design and implementation through testing, deployment, monitoring, and support.
- Use sound engineering judgment to balance performance, cost, scalability, maintainability, and data quality in technical decisions.
- Share knowledge through code reviews, design discussions, documentation, and collaboration with earlier-career engineers.
Minimum qualifications
- 5+ years of data engineering experience building production-grade data solutions and platforms at scale.
- Experience with batch processing technology
- Strong proficiency in SQL, including complex joins, window functions, query tuning, and performance optimization across distributed compute engines.
- Hands-on experience with Apache Spark and the ability to tune and troubleshoot Spark workloads in production.
- Experience with big data and lakehouse technologies such as Spark, Databricks, Trino/Presto, Iceberg or Delta Lake, and orchestration tools such as Airflow.
- Experience designing and maintaining data models for analytics and BI, including star or snowflake schemas, aggregate tables, and semantic or metric layers.
- Familiarity with distributed data systems, including reliability, monitoring, fault tolerance, and cost or performance trade-offs.
- Experience contributing to well-scoped projects within your domain and collaborating across teams to deliver high-quality solutions.
- Strong ownership mindset, clear communication skills, and the ability to work effectively both independently and collaboratively across time zones.
- Comfort using AI-assisted tools in everyday engineering work, with good judgment on correctness, security, and maintainability.
Preferred qualifications
- Experience contributing to production-grade agent skills, MCP tools, or AI-assisted data workflows.
- Experience with metrics stores, semantic layer platforms, or governed analytics platforms.
- Experience with data governance practices including metadata management, cataloging, lineage, access control, and data quality frameworks.
- Experience supporting executive reporting, self-serve analytics, or experimentation use cases on shared data platforms.
- Exposure to API development or service-based data access patterns.
- Interest in applying AI-assisted tools to improve engineering productivity and platform workflows.
Accommodation requests
Expedia Group is committed to providing an inclusive and accessible recruiting experience. If you need an accommodation or adjustment due to a disability during the application or recruiting process, please submit a request at https://expedia.service-now.com/askeg?id=job_accommodation.
About Expedia Group
Expedia Group includes three flagship consumer brands - Expedia, Hotels.com, and Vrbo - along with a leading B2B travel business and travel advertising offerings. Across our brands and business, we help travelers explore the world with confidence and ease.
Important notice
Employment opportunities and job offers at Expedia Group will always come from Expedia Group's Talent Acquisition and hiring teams. Never share sensitive personal information unless you are confident of the recipient. Expedia Group does not extend job offers via email or messaging tools to individuals with whom we have not made prior contact. Our email domain is @expediagroup.com. The official place to find and apply for roles is https://careers.expediagroup.com/jobs/.
Equal Opportunity
Expedia is committed to creating an inclusive work environment with a diverse workforce. All qualified applicants will receive consideration for employment without regard to race, religion, gender, sexual orientation, national origin, disability or age.