Are you passionate about helping people improve the impact of their business through the use of data and analytics? If you enjoy the challenges associated with making complex information systems more user-friendly and effective, we are looking for you to join a team of creative and highly motivated professionals who are driving innovation in the upstream oil field services industry!
At National Oilwell Varco, we strive to lead technology innovation that delivers significant value to our customers. We are hiring a Big Data Engineer to help us build and support our data and analytics delivery pipeline.
Responsibilities
- Design, develop, and optimize scalable data ingestion, transformation, and analytics pipelines for IoT, operational, and business data using Spark and Delta Lake technologies.
- Create and tune Spark transformation, aggregation, streaming, and machine learning workloads, including optimization of Spark clusters and processing infrastructure.
- Develop reusable frameworks for ingesting, validating, enriching, and processing large-scale IoT and event-driven datasets.
- Build and maintain analytical data models, curated datasets, and KPI calculation frameworks that enable product analytics, operational reporting, business intelligence, and self-service analytics.
- Collaborate with Product, Engineering, Operations, and business stakeholders to define, implement, and maintain meaningful KPIs, metrics, and data products that support decision-making.
- Design and implement real-time and batch processing solutions using event-based and streaming technologies.
- Build and execute large-scale data migration solutions for transferring historical and operational data from source systems such as OSI PI and TimescaleDB into enterprise data platforms.
- Develop monitoring, observability, reconciliation, and automation tools that ensure the reliability, quality, and performance of data pipelines, migrations, analytics workloads, and MLOps/DataOps processes.
- Support the delivery of related platform capabilities, including APIs, data services, and integrations, while working within Agile and DevOps methodologies.
- Document solution architecture, data lineage, business logic, KPI definitions, and operational procedures to support long-term maintainability and governance.
Requirements
- Bachelor’s degree in computer science, information systems, or a related field. However, relevant experience will be considered.
- A minimum of 5 years of relevant experience required.
- Advanced understanding with spark and/or Databricks(preferred), including spark cluster tuning.
- Great understanding of distributed systems and partitioning.
- Automation and scripting using .net, C#, Python, Javascript, GoLang, AWS Cloud APIs.
- Database technologies and Timeseries databases like OSIPI, PostgreSQL, Timescale and SQL..
- Developing applications and/or scripts utilizing Timeseries data.
- Linux, Windows OS, Containers.
- Data management and data integration strategies.
- Amazon Web services and/or cloud technologies.
- Visualization/reporting tools such as Tableau, Spotfire, Grafana, etc.
- Understanding of CI/CD practices and technologies. GitHub and GitHub Actions(preferred), Jenkins, TeamCity, Udeploy, etc.
- Optimally communicating technical information.
- Working knowledge of ML and AI model development and the data science development life cycle is an added advantage.
- Experience with large datasets and understanding of stream and batch processing , Datadog, Brokers like Kafka or NATS , Container technologies, and terraform.
- Ability to relate architectural decisions and recommendations to business needs.
- Strong analytical and problem-solving skills.
- Highly independent and ability to communicate across all levels of the organization and work with diverse projects teams.
- Willingness to demonstrate a growth-mindset when faced with new challenges and opportunities.