HeyMilo is building AI interviewers that automate and improve hiring through conversational AI. We work closely with companies to bring AI into real hiring workflows. We also run an applied AI team that studies where our proprietary models succeed and fail.
The Role
We're hiring a Research Engineer to join our Applied AI team in Colombo, focused on HR and recruiting. You'll design and build reinforcement learning environments with verifiable rewards for real workflows across the HR/recruiting stack: sourcing, screening, scheduling, applicant tracking systems, and the systems of record around them. You'll build the simulators, reward functions, and evaluation harnesses that show where models succeed and fail across this stack, and that let us improve them.
The work is equal parts ML research and systems engineering. You'll take a workflow from problem definition to a reproducible environment that models can be evaluated and trained against.
What you'll do
Design and build RL environments for HR and recruiting workflows: realistic simulators of candidates, recruiters, and ATS platforms, tool interfaces, and episodic task generation with proper isolation and reproducibility
Evaluate where models succeed and fail across the HR/recruiting stack, from multi-step recruiter workflows to integrations and record matching across systems
Design verifiable reward functions that score correct intermediate actions as well as end states, and hold up against reward hacking
Build and maintain evaluation harnesses that run task suites across frontier and open-weight models, with clean scoring and cost tracking
Run post-training experiments (e.g. RLVR-style fine-tuning of open models) to validate that your environments produce a learnable signal
Package environments and results for reproducibility, and contribute to research write-ups and published evaluations
Collaborate with analysts who author tasks and rubrics, turning their ground truth into running environments
What we're looking for
PhD in AI, Machine Learning, or a closely related field (required)
Solid grounding in reinforcement learning and LLM post-training (reward design, policy optimization, evaluation methodology)
Strong software engineering skills in Python, with code that others can run and build on
Hands-on experience with LLMs: running evaluations, building agentic loops, tool calling, fine-tuning
Comfortable with containers and infrastructure (Docker, Linux, cloud environments) for reproducible experiment setups
Ability to operate in ambiguity and move quickly; comfortable owning a problem end to end
Bonus
Published research or open-source contributions in ML, RL, or evaluation
Experience with RL/eval frameworks and simulated or sandboxed environments
Experience training or fine-tuning open-weight models at any scale
Familiarity with HR tech: applicant tracking systems, recruiting workflows, or staffing operations
Why join
Ground-floor role on a new applied AI team with real influence over how we build
Work reviewed by and published alongside experienced researchers
High visibility, fast-paced, execution-driven environment
Competitive pay and benefits