Staff AI/Machine Learning Engineer
About Tonic
Tonic AI builds the data infrastructure behind modern AI. Our products de-identify real enterprise data for safe use in training and evaluation, and generate synthetic environments that agents can be trained and tested in. We work with frontier AI labs pushing the edge of what models can do, and with hundreds of enterprises including Fidelity, JPMorgan, and Comcast who need to use their real data safely. This role sits at the center of both.
About The Role
As a Staff AI Engineer at Tonic, you'll own the models that make Tonic's data trustworthy - training the synthesis models that replace sensitive data with something realistic enough to stay useful downstream, and the NER systems that detect it across legal, clinical, and enterprise text. You'll also build the synthetic environments agents train and get evaluated in, and the eval infrastructure that actually separates frontier models instead of another benchmark everyone's saturated.
The stakes are real. Your models run inside customer environments handling genuinely sensitive data, where "close enough" isn't good enough. You'll work both sides of the frontier - partnering directly with the labs training the next generation of models, and enterprise teams trying to ship AI safely.
What You'll Do
Design and build the systems that generate longitudinally coherent synthetic environments for agent training and evaluation, including persona modeling, task generators, and verifiable ground truth.
Build and maintain synthesis models that generate realistic replacement values at very large scale, preserving format, statistical distribution, and semantic consistency so de-identified data stays useful downstream.
Train and improve the NER models behind our entity detection, driving accuracy and recall across free text, structured fields, and mixed enterprise data at scale.
Build evaluation infrastructure that grades agent outcomes, not just traces, and produces real discrimination between frontier models on real tasks.
Fine-tune and evaluate open-weight models on Tonic-generated data, and turn benchmark results into product and research direction.
Expand coverage into new domains, languages, and entity types, and handle the long tail of formats and edge cases that real customer data throws off.
Own model evaluation across the board: precision and recall on detection, utility preservation on synthesis, and outcome-level grading for agents.
Optimize inference so models run efficiently on large volumes of sensitive data inside customer environments.
Partner directly with frontier labs and enterprise ML team to turn hard data problems into shipped model improvements.
Set technical direction for a small, senior team and raise the bar on rigor, reproducibility, and shipping.
What You’ll Bring
8+ years (or PhD with 3+ years) building production ML systems, with real depth in some combination of LLMs, agents, RL, NER, or information extraction.
Hands-on experience training and shipping models to production, and a pragmatic bar for quality: you know how to measure it, where it breaks, and when it's good enough to ship.
Experience with generative or synthesis models where output fidelity and downstream utility both matter, not just plausibility.
Strong software engineering fundamentals. You write code others build on.
Fluency with modern training and eval stacks (PyTorch, distributed training, standard agent and benchmark frameworks).
Comfort working with messy, sensitive, real-world data and the privacy constraints that come with it.
A track record of framing ambiguous problems and driving them to measurable, shipped results.
Bonus: synthetic data generation, data privacy or de-identification, or benchmark construction.
Benefits We Offer
Competitive salary and equity
Unlimited paid time off
401k plan with employer contribution
Medical, dental, and vision insurance
Generous parental leave policy
Remote-friendly work environment