AI Benchmark & Datasets Engineer / Researcher Ajouter aux favoris
- Proactively identify, prioritize, and curate relevant public and client-driven benchmarks across our target use cases and markets.
- Evaluate candidate benchmarks for clarity, data quality, evaluation methodology, and fit with our model roadmap.
- Run benchmarks with baseline models to validate setup, uncover edge cases, and de-risk R&D runs.
- Hand off "benchmark-ready" packages to R&D (specs, data, evaluation scripts, expected metrics, constraints)
- Maintain a shared vocabulary and documentation around benchmarks, datasets, and evaluation formats that GTM and R&D can both use.
- Track and organize benchmark results, model leaderboards, and "what good looks like" for different customers and scenarios.
- Contribute to demos and public-facing proof points based on benchmark outcomes.
Cover letter It's always a pleasure to say hi! If you could leave us 2-3 lines, we'd really appreciate that. You are expected to meet at least one of the following criteria:
- You have published at least one paper at NeurIPS, ICLR, or ICML - where you were the lead author or made significant conceptual & code contributions.
- You have significantly contributed to an LLM training effort which became newsworthy (topped a Hugging Face benchmark, best in class model, etc.), preferably using multiple GPU's.
- You have spent at least 6 months working in a leading Machine Learning research center (e.g. at: Google Brain / Deepmind, Apple, Meta, Anthropic, Nvidia, MILA).
- You were an ICPC World Finalist, or an IOI, IMO, or IPhO medalist in High School.
- Have experience with ML/LLM evaluation , data science, or technical product roles, ideally around benchmarks or experimentation.
- Are comfortable reading papers , leaderboards , and Github repos, and turning them into clear, r epeatable benchmark specs.
- Can talk comfortably with both engineers and customers, and translate between technical detail and business value .
- Care about high-quality data , reproducible experiments, and crisp documentation
- Are respectful of others
- Are fluent in English
- Published or open-sourced work on LLM evaluation, benchmarking or data quality.
- Experience designing custom benchmarks or evaluation protocols for novel model capabilities.
- Join an intellectually stimulating work environment.
- Be a pioneer: you get to work with a new type of "Live AI" challenges around long sequences and changing data.
- Be part of one of an early-stage AI startup that believes in impactful research and foundational changes.
- Type of contract : Full-time, permanent
- Preferable joining date : Immediate. The positions are open until filled - please apply immediately.
- Compensation : based on profile and location.
- Location : Remote work. Possibility to work or meet with other team members in one of our offices: Palo Alto, CA; Paris, France or Wroclaw, Poland. Candidates based anywhere in the EU, UK, United States, and Canada will be considered.
Emplois Recommandés
Développeur Front-End & Data (H/F)
Créée en 2010, Forward Global est une société à mission certifiée B Corp™, rassemblant les expertises de plus de 400 Talents dans le monde, qui œuvrent à réduire les risques numériques, économiques, …
Chargé de marketing opérationnel H/F - CDD
JD SPORTS, QUI SOMMES-NOUS ? JD Sports c’est le Roi Anglais Incontesté de la Basket ! Avec plus de 3 500 boutiques dans près de 40 pays, nous sommes présents en Europe, aux States, en Asie et …
Lead Mobile Engineer - React Native Ajouter aux favoris
Ready to be part of the Legal Tech revolution? As the leading SaaS publisher in Europe, DiliTrust is transforming legal departments worldwide with cutting-edge technology Our Impact: From Gen…
Comptable senior (H/F)
Les missions principales pour le poste Comptable senior (H/F) Nous recherchons pour notre client, cabinet comptable , un(e) Comptable Senior (H/F) en CDI. En ligne direct avec une experte com…
Garde d'enfants, Babysitting (H/F)
Recherche Baby Sitter H/F pour la garde de deux enfants âgés de 1 an et de 6 ans du 01/09/2026 au 30/06/2027 sur les horaires suivants : Lundi, Mardi, Jeudi, Vendredi de 17h00 à 19h00. Toutes nos of…
Customer Success Manager Marchés Publics (H/F - CDI)
Mooncard est la Fintech française qui révolutionne les dépenses professionnelles avec une solution innovante et intelligente. Notre mission? Libérer les professionnels des tâches administratives fast…
FROREUR HORIZONTAL HDD (H/F)
Description du poste : MyConnectt, c'est l'agence de recrutement et d'intérim digitale du Groupe Connectt. Accessible de partout en ligne et disponible 24/24h et 7/7j, nous proposons des postes dans p…
Nounou 6 h/semaine à PARIS(14) pour 3 enfants, 1 an, 4 ans, 6 ans
Description de l'offre: Description de l'offre : Pour une de nos familles, nous sommes à la recherche d'un(e) nounou à domicile à PARIS(14) pour 6 heures de travail par semaine pour garder 3 enf…
PROJETEUR BÉTON ARMÉ (H/F)
Le poste : À propos de nous Bureau d'études structure intervenant sur des projets de logements collectifs, tertiaire et équipements publics, nous accompagnons les maîtres d'ouvrage et architecte…
Gestionnaire Immobilier H/F
Five Guys est une entreprise américaine familiale spécialisée dans le burger « gourmet ». Tous nos produits sont frais et de qualité. Nos clients peuvent eux-mêmes élaborer leurs burgers parmi nos…