Lead AI/ML Data Scientist- Vice president
About the Team:
Citi is looking for a Lead AI/ML Data Scientist to join the Olympus Data Reconciliation and Engineering team, where you will shape the next generation of AI and machine learning capabilities powering enterprise-scale reconciliation across global processing hubs.
In this role, you will drive the full lifecycle of ML model development — from ideation and architecture through to deployment and adoption — delivering measurable impact across Capital Markets operations, risk, and finance. Your work will sit at the intersection of advanced data science and real-world financial systems, influencing outcomes at a global scale.
Responsibilities:
- Design, build, and deploy AI and machine learning models — including Agentic AI and Generative AI solutions — to solve complex reconciliation and data engineering challenges at enterprise scale.
- Lead the end-to-end ML model development lifecycle, from requirements gathering and data preprocessing through to ensemble modeling, validation, and production integration.
- Analyze large volumes of structured and unstructured financial data to uncover trends, patterns, and opportunities for optimization across banking platforms.
- Define and deliver ML model roadmaps in collaboration with technical and business teams, ensuring alignment with project timelines, budgets, and Citi's architecture standards.
- Translate complex data findings into clear visualizations and strategic recommendations that inform decisions made by senior business and technology leaders.
- Partner with engineering, operations, and cross-functional teams to ensure seamless model integration, long-term scalability, and reliable performance in production environments.
- Identify and communicate technology risks and their business implications, developing mitigation strategies and maintaining transparency with stakeholders at all levels.
- Maintain comprehensive model documentation and support knowledge transfer to ensure continuity and adoption across teams.
Required Qualifications & Skills:
Technical Expertise:
- 10+ years hands-on experience in AI/ML development and big data engineering within Financial Services, Insurance, or Telecom environments
- Expert-level proficiency in Python (scikit-learn, TensorFlow, PyTorch, Pandas, NumPy), R (caret, tidyverse, mlr3), and SQL (PostgreSQL, Oracle, MySQL)
- Deep technical knowledge implementing supervised and unsupervised ML algorithms: linear/logistic regression, neural networks (CNN, RNN, LSTM, Transformers), k-means clustering, DBSCAN, decision trees (CART, C4.5), and ensemble methods (Random Forest, XGBoost, LightGBM, CatBoost)
- Proven experience building and deploying Agentic AI and LLM-based solutions using:
- LangGraph for complex agent orchestration and state management
- LangChain for chain-of-thought reasoning and retrieval-augmented generation (RAG)
- Agent Development Kit (ADK) for enterprise-grade autonomous agent development
- Production-level experience with MLOps frameworks and infrastructure:
- Apache Airflow for ML pipeline orchestration and workflow automation
- Kubernetes for containerized model deployment and scaling
- Docker for reproducible ML environments
- Advanced proficiency with distributed computing technologies:
- Apache Spark (PySpark, Spark MLlib) for large-scale data processing
- Hadoop ecosystem (HDFS, MapReduce, YARN)
- Apache Hive for data warehousing and SQL-on-Hadoop
- Expertise with cloud-native data platforms:
- AWS S3 for scalable data lake storage
- Amazon Redshift for enterprise data warehousing
- AWS SageMaker, Azure ML, or Google Vertex AI (beneficial)
- Strong background in data reconciliation frameworks, data quality validation, and ETL/ELT pipelines for financial data processing at enterprise scale
Beneficial Skills & Qualifications:
- Hands-on experience with advanced statistical modeling: Generalized Linear Models (GLM), Random Forest, Gradient Boosting (AdaBoost, XGBoost), and Natural Language Processing (NLP) techniques including text mining, topic modeling (LDA), and sentiment analysis
- Experience with model versioning and experiment tracking tools (Mlflow, Weights & Biases, DVC)
- Proficiency with Git/GitHub/Bitbucket for version control and collaborative development
- Knowledge of CI/CD pipelines for ML model deployment (Jenkins, GitLab CI, GitHub Actions)
- Familiarity with data visualization libraries (Matplotlib, Seaborn, Plotly) and BI tools (Tableau, Power BI)
- Experience with real-time streaming data frameworks (Kafka, Kinesis)
- Passion for staying current with emerging AI/ML frameworks, research papers, and open-source contributions
Education:
Bachelor’s or Master’s degree in Computer Science, Data Science, Software Engineering, Information Systems, Mathematics, Statistics or related fields of study.
\------------------------------------------------------
Job Family Group:
Technology
\------------------------------------------------------
Job Family:
Data Science
\------------------------------------------------------
Time Type:
Full time
\------------------------------------------------------
Most Relevant Skills
Please see the requirements listed above.
\------------------------------------------------------
Other Relevant Skills
For complementary skills, please see above and/or contact the recruiter.
\------------------------------------------------------
Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.
If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.
View Citi’s EEO Policy Statement and the Know Your Rights poster.