Applied AI, Evaluation Engineer

mistral.ai
Paris

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

About Mistral

At Mistral AI, we believe in the power of AI to simplify tasks, save time, and enhance learning and creativity. Our technology is designed to integrate seamlessly into daily working life.

We democratize AI through high-performance, optimized, open-source and cutting-edge models, products and solutions. Our comprehensive AI platform is designed to meet enterprise needs, whether on-premises or in cloud environments. Our offerings include le Chat, the AI assistant for life and work.

We are a dynamic, collaborative team passionate about AI and its potential to transform society.

Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between France, USA, UK, Germany and Singapore. We are creative, low-ego and team-spirited.

Join us to be part of a pioneering company shaping the future of AI. Together, we can make a meaningful impact. See more about our culture on .

About The Job

The Applied AI team is Mistral's customer-facing technical organization. We work directly with enterprise clients from pre-sales through implementation to deploy cutting-edge AI solutions that deliver measurable business impact. Our team combines deep ML expertise with strong customer engagement skills, operating like startup CTOs who own end-to-end project execution.

However, the AI graveyard is full of great ideas nobody could measure or prototypes that never made it to production. As a first Evaluation Engineer , you'll design the methodology, build the infrastructure, and define what "ready for production" means across verticals and use cases.

You will design and implement evaluation systems that help our customers understand model performance across their specific use cases, build robust evaluation infrastructure, and work closely with both research and customer-facing teams.

Research builds evals for frontier capabilities but customers don't care about MMLU scores. We need in Applied AI evals and frameworks for customer reality domain-specific, risk-aware, production-grade . The kind that tell you whether your medical summarization model will hallucinate drug interactions, or whether your legal assistant will invent case citations.

This role sits at the intersection of research, engineering, and solutions , you will play a critical cross role in measuring, understanding, and improving the capabilities of our models for our enterprise customers.

What you will do

- Design and implement comprehensive evaluation frameworks to measure LLM capabilities across diverse customer use cases, including text generation, reasoning, code, and domain-specific applications

- Build scalable evaluation infrastructure and pipelines that enable rapid, reproducible assessment of model performance

- Develop novel evaluation methodologies to assess emerging capabilities or verticalized use cases (cybersecurity, finance, healthcare, etc.) and enable the Solutions (Deployment Strategist and Applied AI) on these topics.

- Create custom evaluation suites tailored to enterprise customers' specific needs, working closely with them to understand their requirements and success criteria

- Collaborate with research teams to translate evaluation insights into model improvements and training decisions

- Partner with product teams to continuously improve our evaluation tooling based on customer feedback

How We Work in Applied AI

- We care about people and outputs.

- What matters is what you ship, not the time you spend on it

- Bureaucracy is where urgency goes to vanish. You talk to whoever you need to talk to. The best idea wins, whether it comes from a principal engineer or someone in their first week.

- Always ask why. The best solutions come from deep understanding, not from copying what worked before

- We say what we mean. Feedback is direct, timely, and given because we care.

- No politics. Low ego, high standards.

- We embrace an unstructured environment and find joy in it.

About you

- You are fluent in English

- 3+ years of experience in ML evaluation, benchmarking for LLM or agentic systems

- You have proven experience in AI or machine learning product implementation with APIs, back-end

- You have deep understanding of concepts and algorithms underlying machine learning and LLMs

- You have strong technical coding skills in Python

- You hold strong communication skills with an ability to explain complex technical concepts in simple terms with technical and non-technical audiences

Ideally you have:

- Contributions to open-source evaluation frameworks (e.g., LM Eval Harness, OpenAI Evals) or published research on LLM evaluation

- Experience as a Customer Engineer, Forward Deployed Engineer, Sales Engineer, Solutions Architect or Technical Product Manager

- Experience with ML frameworks (PyTorch, HuggingFace Transformers)

Benefits

PTO : The CDI contract will be a "Forfait 218 jours", corresponding to 25 days of holidays and on average 8 to 10 days of RTT days, and complete autonomy on working hours

⚕️ Health : Full health insurance coverage for you and your family

Transportation : We offer a €600 annual mobility allowance. This package covers 50% of your public transportation costs and includes the Sustainable Mobility Allowance (FMD), encouraging eco-friendly travel options such as cycling or carpooling.

Food : Swile meal vouchers with 10,83€ per worked day, incl 60% offered by company

Sport : Gymlib - sponsorship by Mistral of a significant part of the monthly fee (depending on the program you chose)

Parental policy : 4 additional weeks for parents on top of what is offered by the French state.

By applying, you agree to our Applicant Privacy Policy .

What we offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page .

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy .

Publié le 2026-06-24

Emplois Recommandés

Consultant débutant (F/H) - Expertise comptable - Paris - septembre 2025

EY
Paris

Vous serez amenés à réaliser des missions telles que : La gestion de travaux de production comptable et reportings financiers Les travaux de clôture de fin d'exercice La conversion de comptes …

Voir les Détails
Publié le 2026-03-15

Senior Mechanical Design Engineer

Interstellar Lab
Paris

Senior Mechanical Design Engineer Full-time In-person   At Interstellar Lab, our mission is to build a future full of life, on Earth and beyond. We believe in a world where technology en…

Voir les Détails
Publié le 2026-06-21

Ingénieur électricité Très Haute Tension (HTB) H/F

Eng'IN Technologies
Paris

Le poste : Notre équipe en pleine expansion, recherche un nouvel Ingénieur électricité Très Haute Tension (HTB) H/F et c'est peut-être vous ! Intervenir sur des projets à forte valeur ajoutée …

Voir les Détails
Publié le 2026-06-18

CADRE DE SANTE (H/F) - INTRA HOSPITALIER UNITE SARAH BERNHARDT - SITE AVRON - 75020 PARIS

Ghu Paris Psychiatrie & Neurosciences
Paris

Uunité Sarah BERNHARDT Lieux de travail 129 rue d’Avron 75020 PARIS VOUS SOUHAITEZ REJOINDRE UN ACTEUR HOSPITALIER MAJEUR DANS LA PRISE EN CHARGE EN PSYCHIATRIE ET NEUROSCIENCES ?  Le Groupe H…

Voir les Détails
Publié le 2025-12-20

Manager logistique (F/H) BASSIN 77

Leroy Merlin France
Paris

Envie de manager une équipe et d’avoir un impact concret, chaque jour ? Le poste de Manager logistique, appelé Cheffe/Chef de secteur logistique, mêle organisation, lien humain et pilotage des flux…

Voir les Détails
Publié le 2026-06-26

Head of Data Analytics Ajouter aux favoris

Paris

Sparteo is an independent suite of AI-powered advertising technologies built on sustainable, sovereign infrastructure. Since 2018, we've been reshaping programmatic advertising to make it faster, …

Voir les Détails
Publié le 2026-06-21

ASSISTANT.E DE DIRECTION - F/H - CDI

Citeo
Paris

Citeo Pro est un éco-organisme dédié aux emballages professionnels. Filiale de Citeo, entreprise à mission engagée de longue date dans la Responsabilité Élargie du Producteur (REP) pour les emballage…

Voir les Détails
Publié le 2026-06-18

Consultant Débutant Treasury en CDI- Tous secteurs - Paris - Septembre 2025 F/H

EY
Paris

Dans le cadre de son développement pour le secteur des Services Financiers, EY recherche des Auditeurs Débutants (F/H) pour rejoindre l'équipe Treasury & Commodities. Au sein du département FSO (Fina…

Voir les Détails
Publié le 2026-03-15

Responsable Territorial.e Seine-Saint-Denis et Nord Pas-de-Calais

Entraide Scolaire Amicale
Paris

L’Entraide Scolaire Amicale (E.S.A) mentore depuis 55 ans des enfants du CP à la Terminale en difficulté scolaire et issus de familles défavorisées. L’objectif est de les conduire vers la réussite et…

Voir les Détails
Publié le 2026-06-18

VENDEUR MATERIEL MONTAGNE / SKI

EKOSPORT
Paris 6e

Tu as envie de partager ta passion pour la montagne et l'outdoor avec nos clients ? Chez Ekosport, chaque journée est l'occasion de conseiller, d'équiper et d'inspirer ceux qui préparent leurs proc…

Voir les Détails
Publié le 2026-06-18