Title: Expert Data Engineer
Shift: 11:00AM - 8:00PM
Work Mode: Hybrid
Location: Bangalore
Introduction to the job:
The AI Data Engineering Lead owns the design, build, governance, and operational readiness of the data pipelines, data products, knowledge assets, and retrieval-ready datasets required to power AI-enabled products.
This role ensures that AI products are not built on fragmented, low-quality, ungoverned, or inaccessible data. The AI Data Engineering Lead works across product, data architecture, AI engineering, platform engineering, security, risk, compliance, MLOps/LLMOps, and operations to ensure that AI products use trusted, permissioned, explainable, and reusable data and knowledge assets.
Some of the things you’ll be doing
1. Own AI-ready data engineering
The AI Data Engineering Lead is accountable for engineering data assets that AI products can safely and effectively use.
- Translate AI product needs into data requirements, data pipelines, data products, and knowledge asset requirements.
- Build and manage data pipelines that support AI-enabled applications, RAG solutions, agents, analytics, automation, and decision-support capabilities.
- Ensure data used by AI products is complete, accurate, timely, traceable, and fit for purpose.
- Create reusable data products that can serve multiple products, business units, and AI use cases.
- Partner with product owners and AI engineers to define what data is required for prompts, retrieval, grounding, classification, extraction, recommendations, and workflow automation.
- Ensure AI data assets are engineered for scale, resilience, security, cost efficiency, and production support.
2. Build trusted data products
The role turns raw data into governed, reusable, product-ready data assets.
- Define and build data products for key domains such as client, entity, product, service, transaction, finance, risk, vendor, employee, jurisdiction, and reference data.
- Establish clear data product ownership, service levels, quality expectations, refresh frequency, and access rules.
- Create data pipelines that support both transactional product needs and AI/analytics needs.
- Define source-of-truth usage and reduce reliance on uncontrolled spreadsheets, local files, and duplicate extracts.
- Ensure data products are documented, discoverable, versioned, and reusable.
- Partner with data stewards to resolve quality, definition, and ownership issues.
3. Engineer data for RAG and knowledge-based AI
The AI Data Engineering Lead ensures documents and knowledge assets can be safely used by AI.
- Prepare policies, procedures, contracts, regulatory content, service playbooks, product documentation, client obligations, and operational knowledge for AI retrieval.
- Define document ingestion, parsing, chunking, embedding, indexing, refresh, and retirement processes.
- Partner with AI engineering to design vector stores, semantic search, retrieval ranking, grounding, and citation patterns.
- Ensure retrieval sources are approved, current, versioned, owned, and access-controlled.
- Validate that AI products retrieve the right content for the right user in the right context.
- Prevent AI products from using obsolete, conflicting, unauthorized, or unapproved knowledge sources.
4. Own data quality and trust controls
AI products amplify data quality issues, so this role establishes trust at the data layer.
- Define data quality rules for critical data elements used by AI products.
- Monitor completeness, accuracy, uniqueness, validity, consistency, and timeliness.
- Build automated quality checks into pipelines.
- Create exception handling, issue management, and remediation workflows.
- Partner with data owners and product teams to prioritize data quality fixes based on business and AI impact.
- Ensure AI products can identify when data is missing, stale, conflicting, or unreliable.
- Provide evidence of data quality for product acceptance, risk review, and operational readiness.
5. Manage metadata, catalog, and lineage
The AI Data Engineering Lead ensures teams know what data exists, what it means, where it came from, and how it is used.
- Capture metadata for datasets, data products, APIs, pipelines, reports, knowledge assets, vector indexes, and AI retrieval sources.
- Ensure data and knowledge assets are cataloged and discoverable.
- Define business and technical metadata needed for AI use.
- Document lineage from source systems through pipelines, transformations, vector stores, prompts, outputs, and consuming products.
- Support auditability by making source-to-output traceability visible where required.
- Partner with data governance teams to ensure definitions, ownership, sensitivity, and usage rules are documented.
6. Embed data access, privacy, and security controls
The role ensures AI data usage respects permissions, sensitivity, and client obligations.
- Implement role-based, attribute-based, jurisdictional, client-specific, and purpose-based access controls.
- Ensure AI products only use data that users, applications, models, and agents are authorized to access.
- Partner with security and privacy teams to classify data sensitivity.
- Prevent sensitive data from being exposed through prompts, logs, embeddings, outputs, or retrieved content.
- Ensure data minimization, retention, masking, encryption, and audit logging requirements are met.
- Design data access patterns that work across applications, APIs, data products, vector stores, and AI agents.
- Support security and privacy reviews for AI-enabled products.
7. Support AI data lifecycle operations
The AI Data Engineering Lead ensures data assets can be operated, monitored, refreshed, and retired.
- Monitor pipeline reliability, latency, freshness, cost, and quality.
- Ensure datasets, embeddings, indexes, and knowledge assets are refreshed on appropriate schedules.
- Version datasets used for training, tuning, evaluation, grounding, and retrieval.
- Support rollback of data, pipeline, embedding, or index changes where necessary.
- Partner with MLOps/LLMOps to manage evaluation datasets, grounding datasets, and prompt/model dependencies.
- Ensure incidents involving data quality, data access, stale knowledge, or retrieval failures can be detected and resolved.
- Support ongoing continuous improvement from user feedback, AI output review, and production monitoring.
What technical skills, experience, and qualifications do you need?
Data engineering capabilities
- Data pipeline engineering
- Data product development
- Data modeling and domain modeling
- ETL/ELT design and orchestration
- API and event-driven data integration
- Data quality rule design and monitoring
- Metadata management and cataloging
- Data lineage and traceability
- Master and reference data awareness
- Data access control and sensitive data handling
- Cloud data platforms and lakehouse/warehouse patterns
AI data capabilities
- Retrieval-augmented generation data preparation
- Document ingestion and knowledge processing
- Chunking, embeddings, vector stores, and semantic retrieval
- Grounding datasets and evaluation datasets
- Dataset versioning for model, prompt, and retrieval evaluation
- Data preparation for AI agents and copilots
- Data freshness, retrieval quality, and source governance
- AI data observability and feedback loops
- Permission-aware retrieval design
- Prevention of sensitive data leakage through AI systems
Business and leadership capabilities
- Ability to translate business and product needs into data requirements
- Strong understanding of data ownership and stewardship
- Ability to explain data quality and lineage issues to business leaders
- Strong collaboration with product, AI, architecture, security, and risk teams
- Practical judgment on what data needs to be centralized, federated, reused, or governed locally
- Strong problem-solving around incomplete, conflicting, or low-quality data
- Ability to build reusable capabilities rather than one-off data extracts