About Cyera
Come join the company building the security operating model for the age of AI. AI has changed how data is used — and security must change with it. Cyera's mission is to empower businesses to accelerate AI Adoption by defining a holistic approach to securing AI - from data to access to model. Instead of perimeter controls and static policies, Cyera provides a unified control plane that understands relationships between data, access, and behaviors across humans, systems, and AI. Backed by the world’s leading investors and working with a large and growing list of Fortune 1000 companies, we are looking for world-class talent to join us as we usher in the new era of data and AI security.
Requirements
About the Role
As a Data Engineer, you'll join the Activity Group — the team behind Access Trail, Cyera's activity monitoring product. Access Trail ingests and processes activity events from a wide range of enterprise data sources at very high throughput, turning billions of raw events into the trusted activity layer that powers investigations, analytics, and Cyera's AI-driven security capabilities. This is production infrastructure that enterprise customers depend on daily, operating at serious scale. You'll work across the full data pipeline — from ingestion and high-throughput Spark and ClickHouse ETLs through data modeling to the analytical layer serving the product — and help re-architect our processing layer as we push past current scale limits. You'll join a high-velocity group shipping weekly to a growing customer base.
What You'll Do
- Partner closely with product and engineering stakeholders to translate strategic requirements into scalable data models and production-ready pipelines.
- Architect and scale high-throughput, multi-stage ETL pipelines that ingest and process billions of activity events, designing incremental processing strategies that handle terabyte-scale datasets while balancing data quality, freshness, performance, and cost.
- Build and optimize distributed processing workloads (Spark) and analytical workloads (ClickHouse), eliminating performance bottlenecks in production environments serving enterprise customers.
- Own data modeling end to end — converting raw activity data from diverse sources into a trusted, well-tested, unified model that powers investigations, analytics, and product experiences.
- Champion data quality and integrity: design validation, testing, and freshness monitoring that catch issues before customers do, and diagnose data quality incidents when they occur.
- Monitor and optimize infrastructure spend across compute, storage, and orchestration, driving measurable efficiency improvements.
Must-Haves
- 4+ years of experience as a data engineer building and operating production data systems.
- Strong proficiency in SQL, including complex transformations, window functions, and performance optimization in columnar/OLAP databases (ClickHouse a strong advantage).
- Hands-on experience with Spark or similar distributed processing frameworks at scale.
- Proven experience designing high-throughput ETL/ELT pipelines, including incremental processing patterns and throughput/cost tradeoffs.
- Strong data modeling skills and a data-quality mindset — building tested, version-controlled transformations.
- Proficiency in Python for pipeline development, tooling, and automation.
- Strong communication skills — able to discuss technical tradeoffs with engineers and translate product needs into data models.
Preferred Qualifications
- Production experience with ClickHouse or other columnar/OLAP systems.
- Experience with streaming and messaging systems (Kafka or similar).
- Experience with AWS data services (S3, EMR/Databricks, MWAA, ECS).
- Familiarity with CDC/data replication tools (Debezium or similar).
- Experience building integrations/connectors with third-party APIs and event sources.
- Experience with multi-tenant data architectures or data access controls.
- Exposure to data security, privacy, or compliance domains.