Mistral AI · Paris/London/Warsaw · Hybrid

Research Engineer, Inference Foundation

10/7/2026

Description

The Role

The Inference Foundation team owns the core of Mistral's inference stack: the inference engine and its orchestration, from the feature set and configuration that serve our models in production to the release machinery that keeps the stack current and production-grade.

This is a hybrid position spanning production LLM serving, engine and platform development, and capacity engineering. You will work on three intertwined problems:

  1. Optimize the inference stack at scale — feature development and fixes in the engine and orchestrator, squeezing more throughput and lower latency out of every GPU under strict quality-of-service targets.

  2. Make capacity elastic — scaling up and down should be cheap and fast, not a performance cliff.

  3. Power the training of our frontier models — high-performance serving that keeps RL and post-training loops running at full speed.

 

What You Will Do

Inference engine & orchestration

  • Develop and fix the core of the inference stack — engine and orchestrator — including feature selection, configuration, and tuning for maximum performance at scale

  • Own the release process for the serving stack: validated, regression-free releases through automated performance gates and progressive rollout

  • Drive improvements and fixes upstream when the open-source engine is the right place for them

Performance & capacity at scale

  • Optimize serving efficiency across the fleet — driving down pod startup time, tackling cold-cache regressions on scale-up, smarter caching and offloading

  • Optimize and maintain the optimal serving topology — overlap communication and transfers with computation, ensure optimal placement, connectivity, and routing

Serving for frontier training

  • Build the serving infrastructure that powers RL and post-training for our frontier models

  • Optimize inference performance across the full spectrum of our workloads

 

Qualifications

  • Experience building and running ML/LLM services at scale, with clear latency and availability targets

  • Hands-on experience with inference engines such as vLLM, SGLang, TensorRT-LLM, or others

  • A solid grasp of inference internals: prefill vs. decode, KV-cache behavior, batching, scheduling, speculative decoding, parallelism strategies

  • Familiarity with distributed and disaggregated serving architectures

  • Comfortable debugging across the full stack — CUDA/NCCL, kernels, containers, networking, storage

  • Python for systems tooling and backend services; PyTorch

  • Kubernetes for running infrastructure at scale

  • GPU and networking fundamentals: CUDA runtime, NCCL, InfiniBand/RDMA

Nice to have

  • Demonstrated vLLM/sglang know-how — upstream contributions, or a track record of running in demanding production environments

  • Hardware-aware optimization for various model architectures

  • Experience serving MoE models at scale (expert parallelism, expert placement/load balancing)

  • CUDA/Triton kernel development; Nsight Systems/Compute profiling

  • Rust and/or C++ in production systems

Application

View listing at origin and apply!

Fast-track your ML job hunt :

Be the first to hear about new sota jobs + exclusive salary research + career cheatsheets.