InterviewHack.ai
Empezar gratis
Vacantes / Prior Labs

ML Engineer, Infrastructure

Prior Labs·Berlinsenior

En corto

  • →Ingeniero de infraestructura para modelos tabulares de IA a escala de producción.
  • →Gestiona clusters GPU multi-proveedor, optimiza costos y mejora el rendimiento de entrenamiento distribuido.
  • →Inversión de 10M+ euros en GPU al año, con impacto directo en decisiones de alto costo.

Fluent technical English (required for collaboration with global teams)

Postularme en la empresa ↗Compartir por WhatsApp

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Gratis, sin tarjeta.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

¿Qué piden?

  • ✓3+ años gestionando infraestructura GPU a gran escala.
  • ✓Experiencia probada con Slurm en entornos multi-tenant.
  • ✓Conocimiento profundo de sistemas: memoria, GPU, comunicación distribuida.
  • ✓Fluidez en Python y PyTorch (internals, profiling).
  • ✓Habilidad para tomar decisiones que mejoren throughput o reduzcan costos.
  • ✓Experiencia con herramientas de productividad de ML (CI/CD, wandb, model registry).

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

SlurmGCPDockerwandbGitHub ActionsuvPyTorchTritonClaude CodeCursor

¿A quién escribirle en Prior Labs?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

Who we are Foundation models transformed text and images. Structured data - the largest and most consequential data format in the world - stayed untouched, until now. What LLMs did for language, we're doing for tables. We pioneered tabular foundation models: TabPFN v2 was a Nature cover story, has passed 3.5M+ downloads and 7,500+ GitHub stars, and runs in production from detecting lung disease with Oxford Cancer Analytics to preventing train failures with Hitachi . The hardest problems - millions of rows, real-time inference, entirely new modalities - are still open, and no one else is working on them at this level. We're a small, highly selective team of 40+ with backgrounds from Google, DeepMind, Meta, Apple, Amazon, Jane Street, and CERN, led by Frank Hutter , Noah Hollmann , and Sauraj Gambhir , and advised by Bernhard Schölkopf and Turing Award winner Yann LeCun. In July 2026, less than 18 months after our €9M pre-seed, we joined SAP as an independent frontier AI lab - same team, mission, and open-weights models, now backed by more than €1 billion over four years. About the Role We spend tens of millions per year on GPU compute to train tabular foundation models. That's not a target, it's what we're running today, and it's growing. The person who owns this infrastructure makes decisions worth millions of dollars: cluster architecture, scheduling efficiency, provider strategy, hardware selection. A wrong call costs six figures. Today we run Slurm on GCP across multiple clusters. We're scaling to multi-cluster, multi-provider infrastructure and evaluating new hardware generations as they come online. You own the full stack, from cluster operations and cost optimization to distributed training performance and the tooling layer that keeps researchers moving fast. You work directly with the research team and understand what they're doing well enough to make infrastructure decisions that actually help them. And this isn't a pure support role. We operate an open environment. If you've got the next SOTA tabular architecture up your sleeve, go ahead and train it. What you'll work on: Own and evolve multi-cluster GPU infrastructure. Slurm on GCP today, multi-provider and new hardware tomorrow. Architecture, scheduling, reliability, cost optimization Drive GPU utilization and training throughput: profiling, memory optimization, communication bottlenecks, systems-level debugging of distributed training across large runs Architect the next generation of our infrastructure: multi-cluster orchestration, new GPU generations, provider diversification, capacity planning against growing compute demands Build the developer productivity layer: CI pipelines, experiment tracking, model registry, data processing, and internal tooling that keeps research iteration speed high Own the compute budget. Tech stack: Slurm, GCP, Docker, wandb, GitHub Actions, uv, PyTorch, Triton You may be a good fit if you have: 3+ years building and operating production GPU infrastructure or distributed training systems at scale. At a major AI lab, a well-funded ML startup, or an HPC environment Deep hands-on experience with Slurm and cluster management. You've debugged scheduling failures, optimized utilization across multi-tenant GPU workloads, and operated infrastructure where downtime has real cost Expert-level systems thinking: memory bandwidth, GPU profiling. You reason about hardware, not configs Strong Python and genuine fluency with PyTorch internals. Enough to profile a training run and tell whether the bottleneck is data loading, communication, or compute Track record of making infrastructure decisions that measurably improved training throughput or cost efficiency Strong AI tooling skills.

Más empleos como este

Empleos remotos de AI / ML Engineer

¿Buscando algo parecido?

Dejá tu email y te avisamos cuando salgan vacantes que coincidan con tu perfil.

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesEmpresas contratandoRevisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Reporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaEs gratis

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY

Vacantes similares activas

Full Stack Engineer, ML Platform

Prior Labs · Berlin

→

Research Scientist Intern (PhD)

Prior Labs · Berlin

→

Research Scientist, Foundational Data Science

Prior Labs · Berlin

→

Technical Recruiter (Berlin)

Prior Labs · Berlin

→