InterviewHack.ai
Empezar gratis
Vacantes / Mistral.ai

Research Engineer - Eval Platform

Mistral.ai·Parissenior

En corto

  • →Construyes infraestructura para evaluar modelos de IA de forma reproducible y escalable.
  • →Trabajas con investigadores para transformar necesidades de evaluación en herramientas confiables y accesibles.
  • →Destaca por garantizar que los resultados de evaluación sean precisos, incluso en modelos agenticos y multi-turno.

Fluency in English is required.

Postularme en la empresa ↗Compartir por WhatsApp
✓ Gratis para empezar✓ Corre en tu navegador✓ Primer dossier sin tarjeta✓ Listo en ~1 minuto

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Tu primer dossier es gratis.

Las preguntas que te van a hacer

1. ¿Cómo garantizarías la reproducibilidad de evaluaciones de LLM cuando cambian las versiones del modelo y el código?

2. Describe cómo diseñarías una API para que researchers accedan fácilmente a resultados de evaluación con baja latencia.

3. ¿Qué métricas o controles usarías para detectar una evaluación ruidosa o sesgada en un entorno de agente multi-turno?

🔒 +7 preguntas más

Sin tarjeta. Subís tu CV y en ~1 minuto tenés el dossier completo.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

💵 USD · Remote · No visa

¿No encontrás lo que buscás? Probá Micro1

Micro1 te ubica directo en empresas de EE. UU. que pagan en USD. Un solo proceso de vetting, múltiples ofertas — sin aplicar en frío.

Que Micro1 te matchee →
📬Vacantes elegidas para TU CV, cada mañana por WhatsApp. Gratis: escribí “vacantes” y el bot te manda tus matches del día. Suscribirme →

¿Qué piden?

  • ✓Master o PhD en Ciencia de la Computación o experiencia equivalente.
  • ✓4+ años construyendo software de producción, especialmente en sistemas distribuidos o ML a gran escala.
  • ✓Excelente dominio de Python y buenas prácticas de diseño de software: pruebas, revisiones, CI/CD.
  • ✓Experiencia corriendo workloads en clusters GPU (Slurm, Kubernetes, Ray, etc).
  • ✓Familiaridad con inferencia y evaluación de LLM.
  • ✓Mentalidad de producto: los investigadores son tus usuarios.

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

PythonSlurmKubernetesRayvLLMSGLangCI/CDAPIsDashboardsLLM inference

¿A quién escribirle en Mistral.ai?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector, co-creating customized AI systems that they can run on their terms. We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited. The Role Evaluation is how we decide which models, checkpoints and recipes ship. As a Research Engineer on the Eval Platform team, you will build the infrastructure every science team relies on to measure model quality, and make it reliable, reproducible and fast. You don't need to have designed benchmarks before. You do need to care about what a score means, and about when a difference between two runs is real. What you will do Build systems that keep eval results reproducible and comparable over time, as models, benchmarks and code evolve. Run evaluations at scale across our GPU clusters, from model serving to scoring. Make eval results easy to access, explore and trust, through APIs and dashboards that researchers use every day. Catch broken or noisy evals before they mislead research decisions. Support evaluation of agentic, multi-turn and tool-using models. Work closely with researchers to turn new evaluation needs into robust, shared tooling. What we're looking for Master's or PhD in Computer Science, or equivalent experience. 4+ years building production-grade software, ideally large-scale ML codebases or distributed systems. Excellent Python and strong software-design instincts: testing, code review, CI/CD. Experience running workloads on GPU clusters (Slurm, Kubernetes, Ray or similar). Familiarity with LLM inference and evaluation. A product mindset: researchers are your users. Self-starter, low-ego, collaborative. Nice to have Experience building or maintaining evaluation harnesses or benchmarks. Hands-on experience with inference engines such as vLLM or SGLang. Experience with agentic or RL environments. Statistics for experimentation: variance estimation, significance testing. Open-source contributions to ML tooling. What We Offer We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks. For the most up-to-date details on benefits available in your location, please refer to our Benefits page . Privacy Policy Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy . Find more English Speaking Jobs in France on Arbeitnow

¿Buscando algo parecido?

Dejá tu email y te avisamos cuando salgan vacantes que coincidan con tu perfil.

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesEmpresas contratandoTodas las herramientasVeredicto de CV (Jev)Cover letter gratisPreguntas de entrevista por rolSimulador de test técnico"Hablame de vos" (respuesta)Revisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Guion de negociación salarialRespuesta STAR gratisTitular + About de LinkedInReporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaPreciosAfiliados — 30%

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY

Vacantes similares activas

HRBP

Mistral.ai · Paris

→

Talent Acquisition Specialist, Early Careers - EMEA

Mistral.ai · Paris

→

Applied AI, Technical Lead, Forward Deployed AI Engineer - London

Mistral.ai · London

→

Applied AI, Forward Deployed Machine Learning Engineer - London

Mistral.ai · London

→