Mistral Cloud - Software Engineer, Managed Kubernetes
About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.
We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.
The Role
We’re looking for passionate and talented Software Engineers to join our Compute team. In this role, you’ll have the opportunity to shape and deliver on-demand Kubernetes clusters with GPU support for Training, Inference, and general-purpose CPU workloads.
More information on Mistral Compute here:Location and Office Policy
We will prioritize candidates who either reside in one of our main offices (Paris, London, NYC) or are open to relocating. We will also consider remote candidates based in the following countries:
EMEA : France, United Kingdom, Germany, Switzerland, Netherlands, Spain, Austria, Poland, Luxembourg
Americas : USA, Canada
We strongly believe in the value of in-person collaboration to foster strong relationships and seamless communication within our team. In any case, we ask all new hires to visit our Paris HQ office (accommodation and travelling covered)
• for the first week of their onboarding
• then at least 3 days every month
What You Will Do
Software Development: Design, build, and maintain scalable control plane services, operators, and custom controllers for Kubernetes
Cluster lifecycle management: Develop automation including provisioning, upgrades, patching, and decommissioning.
Monitoring and Observability: Implement and improve monitoring, alerting, and incident response systems to ensure optimal system performance and minimize downtime
Internal tooling: Create workflows, tools, APIs, and command-line interfaces (CLIs) to empower customers and ML/AI teams to deploy and monitor inference services efficiently.
Reliability: Design resilient systems capable of gracefully handling failures in large-scale distributed environments.
Investigation: identify and resolve unique cluster issues at a low level, documenting findings for future reference.
Incident Response: Participate in on-call rotations to respond to incidents and perform root cause analysis to prevent future occurrences
What We're Looking For
Master’s degree in Computer Science, Engineering or a related field
5+ years of experience in a similar role (Software Engineer, DevOps, SRE ...)
Strong coding proficiency (ideally Golang) and knowledge of software development best practices
Mastery of containerization and orchestration tools (Docker, Kubernetes...)
Exposure to site reliability issues in critical environments (issue root cause analysis, in-production troubleshooting, on-call rotations...)
Experience working against reliability KPIs (observability, alerting, SLAs)
Strong problem-solving abilities and attention to detail.
Ability to own and deliver end-to-end features with minimal oversight.
Excellent communication skills and collaborative attitude.
Team-oriented, humble and eager to learn.
Strong understanding of networking, security, and system administration concepts
Now, it would be ideal if you had experience with:
high-performance computing (HPC) systems and workload managers (Slurm)
monitoring, logging, alerting and observability tools (Prometheus, Grafana, ELK Stack...)
networking, storage, security, and system administration concepts
What We Offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.
For the most up-to-date details on benefits available in your location, please refer to our Benefits page .
Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy .
Emplois Recommandés
AI Engineer (f/m/d)
À propos de Decathlon Decathlon aspire à devenir la meilleure plateforme numérique sportive et l'écosystème ouvert du monde. Nous voulons permettre aux clients de vivre l'expérience Decathlon à tr…
Rénovation, bricolage, maçonnerie à Paris (75009)
Je suis un professionnel de la décoration intérieure. Je peux vous fournir les services suivants: - Peintures (remodelage mural, plâtrage, ponçage, impression). - Carrelage, faïence - Installation …
Comptable Général - Immobilisations, Social, Transport, Fiscalité générale H/F
Dans un contexte de transformation, Polène recrute aujourd’hui un Comptable Général H/F pour prendre en charge les cycles immobilisations et CAPEX, le cycle social et la masse salariale, ainsi que le…
Machine Learning Engineer
About Dashlane Dashlane’s mission is to deliver the credential security every business and employee needs to thrive. Millions of consumers, and over 25,000 brands worldwide, such as Michelin, Air …
Chef.fe de projet analyse de données senior - h/f
Description entreprise : GS1 est l’organisation internationale, neutre et à but non lucratif , créée par les entreprises, pour faciliter et automatiser les échanges entre partenaires commercia…
Head of SEO - M/W
ManoMano est devenu, en moins d’une décennie, un acteur incontournable du secteur de la rénovation et de l’habitat. Créé en France en 2013, ManoMano est la marketplace en ligne de référence pour …
(Stage) Performance Data Analyst
Description Payplug est la solution de paiement française pensée pour les commerçants, e-commerçants de toutes tailles et fintechs. Avec notre plateforme technologique de pointe, nos outils d…
Senior AI Engineer
About Mirakl: Founded in 2012, Mirakl has been at the forefront of marketplace innovation, empowering every business to compete in the platform economy. Today, Mirakl’s operating system combine…
Site Reliability Engineer - SRE
OUR STORY: Join Scaleway and shape the sovereign cloud of tomorrow ! Since 1999, we have been designing secure, sustainable infrastructures aimed at supporting the most ambitious companies. …
Animateur(trice) (H/F)
Le poste de Animateur(trice) (H/F) LA RÉSIDENCE La résidence Beaumarchais est un EHPAD moderne et haut de gamme de 50 places, situé au 63ter Rue Beaumarchais à Montreuil (93). Avec ses équipemen…