
Fast-track your ML job hunt :
As an Evaluation Engineer on the Applied AI team, you will shape how enterprise customers measure and trust AI solutions. You will work closely with clients and internal stakeholders to deliver evaluation systems that define when an AI model is ready for real business impact. The team's mission is to ensure every project—no matter how ambitious—moves from idea to production with clarity and accountability. Your work will sit at the intersection of research, engineering, and solutions, directly influencing how models are improved and deployed across varied domains.
Design and implement comprehensive frameworks to evaluate large language model (LLM) performance across customer use cases
Build and maintain scalable evaluation infrastructure and pipelines for rapid, reproducible assessment
Develop new evaluation methods for sector-specific or emerging model capabilities
Collaborate with customers to create custom evaluation suites that fit their unique needs and criteria
Partner with research teams to turn evaluation insights into actionable model improvements
Work with product teams to refine evaluation tooling based on user feedback
Define and communicate clear success metrics for "production-ready" AI models
At least 3 years of experience evaluating ML, LLM, or agentic systems
Proven background in implementing AI or machine learning products, especially with APIs or back-end systems
Strong Python coding skills
Deep understanding of machine learning concepts and algorithms, especially as they pertain to LLMs
Experience communicating technical topics clearly to both technical and non-technical audiences
Familiarity with evaluation frameworks like LM Eval Harness or OpenAI Evals, or contributions to related open-source projects
Experience working with ML libraries (e.g., PyTorch, HuggingFace Transformers)
Collaborative, direct communicator who values outcomes and low-ego teamwork
Fast-track your ML job hunt :