Senior Machine Learning Engineer, AI Performance
The role
We’re looking for a Senior Machine Learning Engineer to join a high-ownership team responsible for delivering production-ready model releases as our OEM engagements and release cadence accelerate. This is an applied, delivery-focused MLE role—ideal for engineers who love shipping real systems and iterating quickly.
You’ll work on taking models from “works in training” to “meets product constraints,” partnering closely with teams downstream (e.g., inference/performance specialists) to ensure models are ready for deployment on-vehicle. As model capability grows, you’ll help keep the system within tight runtime constraints using a practical model optimisation techniques (e.g., quantisation, distillation, low-rank methods) where appropriate.
Key responsibilities
Own end-to-end delivery of model releases, from initial requirements through training, evaluation, iteration, and final readiness for deployment.
Train and iterate on PyTorch models with a strong experimental approach (hypothesis-driven iteration, ablations, clear evaluation criteria).
Debug and improve model performance using strong analytical skills—identifying regressions, root-causing issues, and proposing fixes.
Apply optimisation techniques (e.g., quantisation and distillation where beneficial), understanding trade-offs and when methods are appropriate.
Collaborate cross-functionally with adjacent ML and performance engineering teams to hand off models, define bottlenecks, and align on optimisation priorities.
Communicate clearly with stakeholders to align on delivery timelines, trade-offs, and readiness criteria.
About you
In order to set you up for success in this role at Wayve, we’re looking for the following skills and experience:
Essential
Proven experience improving performance in production systems with tight constraints (latency, memory, bandwidth, power/thermal, or cost).
Strong hands-on experience training and iterating on deep learning models in PyTorch (not just using high-level tooling).
Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL) and confidence learning adjacent frameworks quickly.
Comfort operating at multiple levels of abstraction — from high-level model behaviour down to low-level kernel/runtime execution.
Familiarity with model optimisation concepts such as quantisation and/or distillation (hands-on is a strong signal, but not a strict requirement if the fundamentals are solid).
Ability to reason across multiple levels of abstraction—from high-level model behaviour down to practical runtime/latency implications.
Strong engineering fundamentals and collaboration skills.
Desirable
Experience working on models that must meet tight latency / efficiency constraints (edge, embedded, real-time, or similarly constrained production settings).
Exposure to ML systems spanning training → evaluation → deployment handoff (even if you’re not writing kernels day-to-day).
Exposure to embedded or edge deployment of ML models, including benchmarking on real devices and handling system-level constraints.
#LI-HH1