
Fast-track your ML job hunt :
The Inference Foundation team owns the core of Mistral's inference stack: the inference engine and its orchestration, from the feature set and configuration that serve our models in production to the release machinery that keeps the stack current and production-grade.
This is a hybrid position spanning production LLM serving, engine and platform development, and capacity engineering. You will work on three intertwined problems:
Optimize the inference stack at scale — feature development and fixes in the engine and orchestrator, squeezing more throughput and lower latency out of every GPU under strict quality-of-service targets.
Make capacity elastic — scaling up and down should be cheap and fast, not a performance cliff.
Power the training of our frontier models — high-performance serving that keeps RL and post-training loops running at full speed.
What You Will Do
Develop and fix the core of the inference stack — engine and orchestrator — including feature selection, configuration, and tuning for maximum performance at scale
Own the release process for the serving stack: validated, regression-free releases through automated performance gates and progressive rollout
Drive improvements and fixes upstream when the open-source engine is the right place for them
Optimize serving efficiency across the fleet — driving down pod startup time, tackling cold-cache regressions on scale-up, smarter caching and offloading
Optimize and maintain the optimal serving topology — overlap communication and transfers with computation, ensure optimal placement, connectivity, and routing
Build the serving infrastructure that powers RL and post-training for our frontier models
Optimize inference performance across the full spectrum of our workloads
Experience building and running ML/LLM services at scale, with clear latency and availability targets
Hands-on experience with inference engines such as vLLM, SGLang, TensorRT-LLM, or others
A solid grasp of inference internals: prefill vs. decode, KV-cache behavior, batching, scheduling, speculative decoding, parallelism strategies
Familiarity with distributed and disaggregated serving architectures
Comfortable debugging across the full stack — CUDA/NCCL, kernels, containers, networking, storage
Python for systems tooling and backend services; PyTorch
Kubernetes for running infrastructure at scale
GPU and networking fundamentals: CUDA runtime, NCCL, InfiniBand/RDMA
Demonstrated vLLM/sglang know-how — upstream contributions, or a track record of running in demanding production environments
Hardware-aware optimization for various model architectures
Experience serving MoE models at scale (expert parallelism, expert placement/load balancing)
CUDA/Triton kernel development; Nsight Systems/Compute profiling
Rust and/or C++ in production systems
Fast-track your ML job hunt :