*Telecommuting role to be performed anywhere in the U.S.
Architect and implement complex, high-volume data pipelines between Snowflake and Databricks utilizing PySpark for distributed data processing and dbt for SQL-based data transformation, including developing macro-driven data quality tests and validation frameworks.
What You Will Do:
- Orchestrate pipeline scheduling, dependency management, and automated failure recovery using Apache Airflow to deliver B2B marketing attribution and multi-touch targeting analytics.
- Administer the enterprise Databricks platform by configuring IAM roles for secure Amazon S3 bucket access, managing application credentials and secrets through Databricks' built-in vault system and OpenShift secrets, and establishing workspace governance policies and cluster configurations for cross-functional data science and engineering teams.
- Design and deploy intelligent retrieval architecture and AI-driven workflows using vector-based search methods and enterprise data platforms, building marketing retrieval and decision-automation applications that integrate multiple data sources and APIs.
- Operationalize MLOps methodologies using MLflow for experiment tracking and model registry management, and Lakehouse monitoring for automated post-production model performance tracking to optimize predictive accuracy and increase marketing return on investment.
- Implement end-to-end machine learning models and deliver stakeholder-facing analytical outputs by building Named Entity Recognition (NER) systems using TextBlob, gensim, and fastText for enterprise systems analysis, developing predictive models using XGBoost and Scikit-learn, constructing deep learning architectures using Keras, and designing time-series forecasting models for event-based user adoption prediction.
- Manage CI/CD pipelines using Git and Tekton to ensure reliable, repeatable code delivery for production applications.
- Build and manage container images using buildah and skopeo, pushing to internal container registries for deployment.
- Lead the deployment and maintenance of containerized data science models and enterprise applications on Red Hat OpenShift (Kubernetes), managing network routes, TLS termination, and container orchestration for highavailability services.
- Lead application security initiatives by completing comprehensive enterprise security compliance assessments encompassing 20+ security controls across the full technology stack, aligned with industry frameworks such as NIST and CIS Controls.
- Perform static application security testing (SAST) using SonarQube, execute vulnerability scanning using Qualys and pip-audit, complete Privacy Impact Assessments (PIA), and conduct STRIDE-based threat modeling.
- Collaborate with enterprise information security teams to remediate identified vulnerabilities, navigate compliance audits, and maintain centralized logging and monitoring through Splunk.
What You Will Bring:
- Master's degree (U.S. or foreign equivalent) in Computer Science or related field and three (3) years of experience in the job offered or related role OR Bachelor's degree (U.S. or foreign equivalent) in Computer Science or related field and five (5) years of experience in the job offered or related role.
- Must have three (3) years of experience with: architecting and implementing high-volume data pipelines between cloud data warehouse (Snowflake) and lakehouse (Databricks) platforms using PySpark for distributed data processing and dbt for SQL-based data transformation, including developing macro-driven data quality test frameworks and validation logic; orchestrating and scheduling data pipeline workflows using Apache Airflow, including configuring DAG-based dependency management, automated failure recovery, and pipeline monitoring for enterprise analytics workloads; administering enterprise Databricks environments, including configuring IAM roles for secure cloud object storage (Amazon S3) access, managing application secrets through platform vault systems and OpenShift secrets, and establishing workspace governance and cluster policies for cross-functional teams; implementing end-to-end machine learning models by: 1) building Named Entity Recognition (NER) systems using TextBlob, gensim, and fastText for enterprise text analysis; 2) developing predictive models using gradient boosting frameworks (XGBoost) and Scikit-learn; 3) constructing deep learning architectures using Keras; and 4) designing time-series forecasting models for event-based prediction; delivering full-scale information retrieval systems for enterprise data by researching, evaluating, and implementing Transformer architectures and Transfer Learning methodologies using deep learning frameworks for semantic search, text classification, and vector-based clustering; operationalizing MLOps methodologies using MLflow for experiment tracking and model registry management, and implementing automated post-production model monitoring to track performance degradation and optimize predictive accuracy; managing CI/CD pipelines using Git and Tekton, building and publishing container images using buildah and skopeo to internal container registries, and deploying containerized applications on Red Hat OpenShift (Kubernetes) with network route management, TLS termination, and high-availability configurations; and leading enterprise security compliance assessments, including performing static application security testing (SAST) using SonarQube, executing vulnerability scans using Qualys, completing Privacy Impact Assessments (PIA), and conducting STRIDE-based threat modeling.
#LI-DNI
The salary range for this position is $158,309 - $180,000/year. Actual offer will be based on your qualifications.
Pay Transparency
Red Hat determines compensation based on several factors including but not limited to job location, experience, applicable skills and training, external market value, and internal pay equity. Annual salary is one component of Red Hat’s compensation package. This position may also be eligible for bonus, commission, and/or equity. For positions with Remote-US locations, the actual salary range for the position may differ based on location but will be commensurate with job duties and relevant work experience.
About Red Hat
Red Hat is the world’s leading provider of enterprise open source software solutions, using a community-powered approach to deliver high-performing Linux, cloud, container, and Kubernetes technologies. Spread across 40+ countries, our associates work flexibly across work environments, from in-office, to office-flex, to fully remote, depending on the requirements of their role. Red Hatters are encouraged to bring their best ideas, no matter their title or tenure. We're a leader in open source because of our open and inclusive environment. We hire creative, passionate people ready to contribute their ideas, help solve complex problems, and make an impact.
Inclusion at Red Hat
Red Hat’s culture is built on the open source principles of transparency, collaboration, and inclusion, where the best ideas can come from anywhere and anyone. When this is realized, it empowers people from different backgrounds, perspectives, and experiences to come together to share ideas, challenge the status quo, and drive innovation. Our aspiration is that everyone experiences this culture with equal opportunity and access, and that all voices are not only heard but also celebrated. We hope you will join our celebration, and we welcome and encourage applicants from all the beautiful dimensions that compose our global village.
Equal Opportunity Policy (EEO)
Red Hat is proud to be an equal opportunity workplace and an affirmative action employer. We review applications for employment without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, citizenship, age, veteran status, genetic information, physical or mental disability, medical condition, marital status, or any other basis prohibited by law.