
Fast-track your ML job hunt :
This role focuses on building and operating the end-to-end execution, training, and data infrastructure that powers Mistral’s agentic models and coding assistants. You will be a core contributor to our agent research stack: designing scalable systems for synthetic data generation, building ultra-fast training and RL execution environments, and maintaining high-throughput execution engines.
You will tackle the engineering challenges at every step of the agent lifecycle: from orchestrating 1M+ concurrent and short-lived sandboxes for untrusted code execution to optimizing agent training codebases, distributed trajectory collection pipelines, and dataset processing workflows across massive hybrid and multi-cloud clusters.
Large-Scale Sandboxing Infrastructure: Design, deploy, and operate our high-throughput sandboxing platform, executing LLM-generated untrusted code across over 1 million isolated environments concurrently for model evaluation and interactive RL environments.
Agent Data Generation Pipelines: Architect and scale high-throughput pipelines for synthetic code generation, agent trajectories, rollouts, and self-play data collection to power post-training and RL loops.
Training Codebase & Systems Optimization: Optimize agent training codebases and distributed execution runtimes (PyTorch, Ray, SLURM/Kubernetes) to minimize multi-step rollout overhead, improve GPU utilization, and eliminate scaling bottlenecks.
Low-Latency Orchestration & Warm Pooling: Reduce sandbox cold-start times to sub-second levels using container warm pools, snapshot/restore technology (e.g., CRIU, microVMs), and optimized image delivery layers across hybrid clusters.
Multi-Cluster Queueing & Resource Allocation: Implement Kubernetes-native custom controllers, CRDs, and queuing systems to dynamically route short-lived evaluation, synthetic data, and agent execution tasks across diverse hardware fleets.
Isolation, Security & Security Boundary: Ensure strict multi-tenant network and process isolation for untrusted agent code using container/sandboxing runtimes (e.g., gVisor, Firecracker) and default-deny network postures.
Operational Excellence: Maintain high availability, telemetry, and automated self-healing across millions of transient jobs while participating in on-call rotations for critical agent training and execution pipelines.
4+ years of experience in Systems Engineering, Distributed Systems, Cloud Infrastructure, or MLOps supporting LLM/RL workloads.
Data & Pipeline Engineering: Proven experience building high-throughput data processing and generation pipelines for large-scale datasets (e.g., Ray, Spark, custom distributed queues).
Deep experience with Kubernetes & Container Tech: Strong expertise writing custom K8s operators/controllers, managing Linux cgroups/namespaces, and optimizing Docker image layers and distribution systems.
High-Performance Software Engineering: Advanced proficiency in Python, Go, C++ or Rust, with a track record of profiling and optimizing high-performance ML or backend systems codebases.
Sandboxing & Isolation Technologies: Hands-on experience with lightweight virtualization, container runtimes, or WASM (e.g., Docker, gVisor, Firecracker).
Queueing & Scheduling: Deep familiarity with task queue systems, resource schedulers, and low-latency queuing architectures for high-volume, short-lived workloads.
Comfort with Ambiguity: Passion for working directly alongside AI researchers to rapidly turn frontier agent ideas into scalable, production-grade infrastructure.
Fast-track your ML job hunt :