InterviewHack.ai
Empezar gratis
Vacantes / Nebius

Senior ML Engineer (Token Factory)

Nebius · Berlinsenior

En corto

  • ▸Ingeniero de ML sénior enfocado en optimizar inferencia y fine-tuning de modelos grandes a escala en GPU.
  • ▸Días a día: perfilar trabajo en GPU, diseñar pipelines de baja precisión y mejorar motores de inferencia como vLLM o TensorRT-LLM.
  • ▸Destacado: trabajar en una de las mayores nubes de GPU del mundo, con miles de GPUs y un impacto directo en el rendimiento de modelos como GPT-OSS o GLM-5.

Excellent command of the English language, alongside superior writing, articulation, and c

Postularme en la empresa ↗Compartir por WhatsApp

En ~1 minuto te damos: quién te entrevista, las preguntas probables con respuestas desde tu CV, y tu CV adaptado a esta vacante. Gratis, sin tarjeta.

🎧¿Llegás a la entrevista? Llevá el copiloto. Nuestra extensión escucha la entrevista en vivo y te muestra anclas de 3-4 palabras desde tu CV y tu preparación — mirás, conectás, hablás. Gratis. Ver la extensión →

¿Qué piden?

  • ✓Conocimiento profundo de arquitecturas de transformadores y ML teórico.
  • ✓Experiencia con herramientas de perfilado de GPU (Nsight, PyTorch profiler).
  • ✓Entendimiento de jerarquía de memoria y tradeoffs de compute/memory en GPU.
  • ✓Familiaridad con conceptos clave de LLM (MHA, RoPE, KV-cache, Flash Attention, cuantización).
  • ✓Experiencia con frameworks modernos de deep learning y alta habilidad en ingeniería de software (Python, CI/CD, testing).
  • ✓Habilidades comunicativas y de liderazgo sólidas.

¿No cumplís todo? Es lo normal — tu dossier gratis te dice qué gaps tenés y cómo cubrirlos en la entrevista.

PythonPyTorchNsightvLLMSGLangTensorRT-LLMTritonCuteCUTLASSCUDA

¿A quién escribirle en Nebius?

Tu dossier gratis identifica a las personas que te entrevistarían — con su background, qué valoran y cómo escribirles para destacar antes de aplicar.

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role Token Factory is a part of Nebius Cloud, one of the world's largest GPU clouds, running tens of thousands of GPUs. We are building a high-performance inference and fine-tuning platform designed to push foundation models to their hardware limits. Our mission is to maximize throughput, minimise latency, and optimise cost-per-token across tens of thousands of GPUs. Some directions we are currently working on, and which you can be a part of: • Inference Optimization: Identifying LLM inference bottlenecks to drive production speedups. Squeezing the maximum performance for a wide range of LLM architectures at scale (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-5). • Inference engines support: Implement novel speculative decoding architectures, optimise components of various LLM designs (dense/MoE, autoregressive/parallel), and contribute to open-source inference engines. • Low Precision Training & Inference: Design and productionise low-precision (FP8, NVFP4/MXFP4) training and inference pipelines with measurable gains in throughput and cost-efficiency. We expect you to have: • A profound understanding of theoretical foundations of machine learning and transformer architecture. • Experience profiling GPU workloads using Nsight, PyTorch profiler, or similar tools • Understanding of GPU memory hierarchy and compute/memory tradeoffs • Familiarity with important ideas in LLM space, such as MHA, RoPE, KV-cache, Flash Attention, and quantisation • Understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.) • Strong software engineering skills (we mostly use Python) • Deep experience with modern deep learning frameworks • Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing • Strong communication and leadership abilities Nice to have: • Experience working with open-source inference engines (vLLM, SGLang, TensorRT-LLM), including contributions • Experience with kernel languages or DSLs such as Triton, Cute, CUTLASS, CUDA • A track record of building and delivering products (not necessarily ML-related) in a dynamic startup-like environment. • Strong engineering skills, including experience in developing large distributed systems or high-load web services. • Open-source projects that showcase your engineering prowess • Excellent command of the English language, alongside superior writing, articulation, and communication skills. Benefits & Perks: • Competitive compensation • Career growth and learning opportunities • Flexibility and ownership • Collaborative and innovative culture • Opportunity to work on impactful AI projects •

No apliques sin prepararte

Investigamos quién te entrevista, adaptamos tu CV y te ensayamos en vivo — gratis la primera.

InterviewHack.ai

Preparate para la entrevista exacta: quién te entrevista, tu CV a medida y coach real.

Producto

VacantesRevisar CV (ATS) gratis¿Cómo suena tu inglés?¿Te pagan bien?Reporte de sueldos LATAMCursos gratisBlogCV a medidaPráctica habladaEs gratis

Empleos remotos

ReactPythonFull-StackLATAMArgentinaMéxicoVer todas →

Preparate

Práctica habladaFrontendBackendAI EngineerPor empresaVendete con tu CV

Empresa

Buscás talentoAcerca deContactoPrivacidadTérminos

© 2026 InterviewHack.ai · Tu CV es tuyo. Nunca se usa para entrenar nada. · Un producto de IA-PTY

Vacantes similares activas

Senior Software Engineer (Token Factory)

Nebius

→

Key Customers Solutions Architect EMEA

Nebius

→

Product Manager - Security

Nebius · Berlin

→

Senior Backend Engineer

Nebius

→