
Fast-track your ML job hunt :
Your mission, alongside the team, is to build and operate the framework to ensure Mistral’s solution delivery is reliable and sustainable - and applied uniformly across all our accounts, both Mistral-hosted and customer-hosted. You should already have a strong understanding of what operational excellence looks like, and you’re ready to scale your impact.
You will operate in four concurrent modes:
- BUILD - Design for a fleet of Mistral platforms and apps. Build proactivity to reduce reactivity.
Productize reliability, author runbooks, create SLO templates, implement observability.
- RUN - Operate the Tier-1 customer environments that Mistral are contracted to operate.
Ensure SLO compliance, own on-call and incident response, manage drift, partner with Technical Support as L3 escalation, champion high signal post-mortems.
- ENABLE - Productize how Mistral deploy, secure, and scale our Applied AI solutions.
Engineer on-demand provisioning, author security baseline packages, embed security guardrails, automate everything.
- SECURE - Own the security operations layer for our customer-side deployments.
Lead CVE response across the fleet, ship supply-chain integrity controls (SBOM, signed images, provenance), co-page with InfoSec on security incidents, enforce secure-config baselines.
This is a framework-first, fleet management role at heart. If you're excited by the difference between solving one customer's problem and structurally solving the class of problem for every customer, this is the role.
• We care about people and outputs.
• What matters is what you ship, not the time you spend on it
• Bureaucracy is where urgency goes to vanish. You talk to whoever you need to talk to. The best idea wins, whether it comes from a principal engineer or someone in their first week.
• Always ask why. The best solutions come from deep understanding, not from copying what worked before
• We say what we mean. Feedback is direct, timely, and given because we care.
• No politics. Low ego, high standards.
• We embrace an unstructured environment and find joy in it.
• Fluent in English.
• 5+ years in SRE, Production Engineering, or DevOps, with a record of shipping tooling.
• Strong multi-tenant Kubernetes fluency, namespace segmentation, network policy, RBAC, admission control, operations at scale.
• On-call discipline: incident response, blameless post-mortem culture, runbook-first mindset.
• Observability stack in production: Prometheus, Grafana, OpenTelemetry, Loki, Tempo, Signoz.
• Infrastructure as code: Terraform, Ansible (or close equivalents).
• Proficient in Python and/or Golang for tooling and automation.
• Security mindset: you treat secure-SDLC, CVE response, and supply-chain integrity as reliability properties of the shipped artifact, not as someone else's job.
• Strong written communication skills: runbooks, post-mortems, and customer-facing incident comms are core deliverables of this role.
• Comfortable operating with high autonomy in an ambiguous, fast-paced environment — and disciplined enough to defend the team's scope when work tries to spill in.
• Solid Linux internals, networking debug, and distributed-systems fundamentals.
• Cloud or application security background (AppSec, K8s security, supply chain — SBOM, cosign, SLSA). At least one of our early hires must bring this; if it's you, flag it.
• Experience operating LLM / model-serving stacks in production
• Experience with multi-cloud or on-prem hybrid customer environments (AWS, GCP, Azure, sovereign clouds).
• Open-source contributions, particularly in SRE, observability, or security tooling.
Fast-track your ML job hunt :