Description
The DUE ML Core London team builds and operates scalable machine learning systems, simulation workflows, and insight tools designed to improve the evaluation and developer onboarding journeys. By combining expert human judgment with advanced machine learning models, we deliver training and evaluation data for hundreds of metrics and components that comprise the Waymo Driver.
We are looking for researchers and software engineers passionate about developing ML techniques for evaluation systems and driving performance improvements across our technology stack.
You will:
- Build scalable systems for training and fine-tuning large-scale generative models to produce realistic and evaluate interesting driving behaviors.
- Lead the implementation, and iteration of novel RL algorithms, reward functions, and training paradigms tailored for generating high-fidelity and insightful driving behaviors
- Lead the development of cutting-edge Deep Learning models and Generative AI (LLM/VLM) solutions to enhance human-led triaging, introduce automation for high-volume workflows, and perform nuanced analysis of self-driving behavior to detect critical anomalies.
- Oversee the production and optimization of machine learning models aiming to assess Waymo’s expansive fleet of vehicles that cumulatively travel millions of miles.
- Proactively monitor and assimilate best practices from within Alphabet and the broader industry to develop a novel Reinforcement Learning from Human Preference (RLHF) based data collection and evaluation system.
- Collaborate closely with multiple teams (e.g., Prediction, Planning, Research), other technical leads, and senior leaderships across Waymo to deliver on key strategic efforts.